04QuestionsCI/CD system

04 · Worked prompt

CI/CD system

Design workflow DAGs, isolated runners, artifacts, logs, retries, and fair scheduling for bursty builds.
40 minInterview blueprint
INTERVIEW RUBRIC

What a passing answer must show

100 points · 45 minutes

  1. 20pts

    Scope the problem

    0–5 min

    Prioritize the core flows, state the scale, and name the non-goals.

  2. 15pts

    Define contracts

    5–10 min

    Identify durable entities, APIs, idempotency, and the source of truth.

  3. 30pts

    Complete the diagram

    10–25 min

    Trace one write path and one read path. Label the commit boundary and async work.

  4. 20pts

    Lead one deep dive

    25–38 min

    Choose the highest-risk trade-off and explain the mechanism, alternative, and cost.

  5. 15pts

    Prove reliability

    38–45 min

    Walk a failure, recovery, metric, bottleneck, and evolution path.

DRAW THIS FIRST

One complete box-and-arrow design

Design workflow DAGs, isolated runners, artifacts, logs, retries, and fair scheduling for bursty builds.

CI/CD system · system architecture
CI/CD system system architecture. The workflow database owns DAG state; runners are leased, isolated, and disposable. Request path: Developer + Git to Webhook/API edge to Workflow orchestrator to Workflow DB. Asynchronous path: Priority task queues to Runner manager. Read path: Run status API to Log/artifact index. External dependency: Sandbox runner pool. Delivery path: Live log gateway.

Write

Developer + Git enters through Webhook/API edge. Workflow orchestrator owns validation and commits the durable record to Workflow DB.

Propagate

Priority task queues separates the committed write from background work. Runner manager can retry safely while it builds Log/artifact index.

Read

Run status API serves from Log/artifact index, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Sandbox runner pool as an explicit dependency.

Say this first: The workflow database owns DAG state; runners are leased, isolated, and disposable.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

The workflow database owns DAG state; runners are leased, isolated, and disposable.

INTERVIEW CONTRACT

Design workflow DAGs, isolated runners, artifacts, logs, retries, and fair scheduling for bursty builds.

CAPACITY QUESTIONS TO QUANTIFY

Bursty pushes · thousands concurrent jobs · minute-to-hour runtimes · untrusted code. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Developer + Git → Webhook/API edge

    Commit, webhook, cancel enters over HTTPS / RPC. Webhook/API edge handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Webhook/API edge → Workflow orchestrator → Workflow DB

    Workflow orchestrator receives the verified webhook, checks invariants and retry identity, then uses state transition to update Workflow DB. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Workflow orchestrator → Priority task queues → Runner manager → Log/artifact index

    Workflow orchestrator emits publish after commit; Runner manager uses consume and project / update to build Log/artifact index. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Webhook/API edge → Run status API → Log/artifact index / Workflow DB

    Run status API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    Runner manager → Sandbox runner pool

    lease runner crosses into Sandbox runner pool. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

  6. 06
    Deliver without changing the source of truth
    Runner manager → Live log gateway → Developer + Git

    Runner manager uses fan-out; Live log gateway returns updates over SSE logs. Sequence IDs, reconnect cursors, and backpressure make delivery resumable without turning a socket into durable state.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Webhook/API edgeVerify, auth, quotaIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Workflow orchestratorExpand DAG + release tasksWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Workflow DBRun + task stateAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Priority task queuesResource-class partitionsDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Runner managerFenced leases + retriesReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Log/artifact indexCursors + signed linksRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Run status APIDAG, attempts, artifactsRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Sandbox runner poolCompile, test, buildExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
Live log gatewayTail task outputConnection and delivery stateSeparates open connections and fan-out pressure from durable domain state.Reconnect, ordering, slow consumers
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
PostgreSQL owns workflow/DAG and attempt state, S3 owns logs/artifacts, and Kafka/SQS or a durable priority queue owns runnable work.
Partitioning / sharding
Partition metadata by organization_id/repository_id and queue by resource class. Keep one run's DAG local; key objects by org/run/task/digest.
Indexes
Unique provider_delivery_id; (run_id, state), (task_id, attempt), (queue_class, priority, ready_at), and artifact digest indexes.
Replication + consistency
DAG transitions and accepted terminal attempts use transactions or conditional writes. Logs and artifact indexes may converge asynchronously.
Cache, queue + recovery
Use an outbox to avoid DB/queue dual writes; runners use fenced leases, heartbeats, bounded retries, and dead letters.
Capacity math
Estimate webhook bursts, concurrent tasks, log MB/sec, artifact GB/day, runner cold start, and queue age per resource class.
Alternative rejected
Keeping the DAG only in a broker makes repair opaque; durable relational workflow state plus queues for ready work is recoverable.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

DAG scheduling

Maintain remaining-dependency counts and atomically release a node once

Scanning the whole graph after every task is wasteful and race-prone
Runner isolation

Use ephemeral microVMs, outbound network policy, short-lived credentials, and immutable base images

The user workload is hostile by default
Retry semantics

Create a new attempt under the same logical task and publish artifacts only after terminal commit

Retries must not overwrite evidence from earlier attempts
05

Failure pressure test

Show detection, containment, recovery, and evidence.

Runner heartbeat lost

Expire the lease and enqueue a new attempt after a fencing check

expired lease count and duplicate runner activity
Webhook replay

Deduplicate provider delivery ID and repo commit tuple

webhook dedupe rate
Log backpressure

Spool locally with bounded buffers and preserve job execution

dropped log bytes and stream lag
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution
SAY THIS WHILE YOU DRAW

A four-part talk track

  1. Scope

    “I’ll prioritize trigger workflows from signed repository events and execute dag stages with retries and cancellation.”

  2. Scale

    “The design changes around bursty pushes · thousands concurrent jobs · minute-to-hour runtimes · untrusted code.”

  3. Decision

    “Ephemeral runners maximize isolation; warm pools reduce startup time at greater cost and risk.”

  4. Risk

    “The first failure I want to pressure-test is: Duplicate webhooks and lost heartbeats can execute deployments twice or leak capacity.”

Reference details

Open these only after you can explain the diagram above without reading.

01Requirements and state lifecycle4 requirements
  • Trigger workflows from signed repository events
  • Execute DAG stages with retries and cancellation
  • Stream logs and preserve artifacts
  • Isolate tenants and support bursty workloads
CI/CD system · state lifecycle
02Data model and APIs4 entities · 3 interfaces

Core entities

WorkflowDefinitionrepo_id, revision, DAG, policyOwner: Workflow service
WorkflowRunrun_id, commit_sha, state, versionOwner: Orchestrator
TaskAttempttask_id, attempt, lease, resultOwner: Scheduler
Artifactartifact_id, digest, producer, retentionOwner: Artifact service

External interfaces

POST /v1/hooks/git

Verify and deduplicate a repository event

POST /v1/runs/{id}/cancel

Version-check and cancel runnable work

GET /v1/runs/{id}/events

Stream logs and state changes from a cursor

03Deep dives and trade-offsChoose one

DAG scheduling

Maintain remaining-dependency counts and atomically release a node once

Scanning the whole graph after every task is wasteful and race-prone

Runner isolation

Use ephemeral microVMs, outbound network policy, short-lived credentials, and immutable base images

The user workload is hostile by default

Retry semantics

Create a new attempt under the same logical task and publish artifacts only after terminal commit

Retries must not overwrite evidence from earlier attempts
04Failures, recovery, and evidence3 scenarios

Runner heartbeat lost

Expire the lease and enqueue a new attempt after a fencing check

expired lease count and duplicate runner activity

Webhook replay

Deduplicate provider delivery ID and repo commit tuple

webhook dedupe rate

Log backpressure

Spool locally with bounded buffers and preserve job execution

dropped log bytes and stream lag
05What makes the answer seniorInterviewer signals
  • Draw the control plane separately from runner data paths
  • The hardest invariant is one terminal task result despite at-least-once execution
  • Staff-level answers discuss fairness, secrets, and fleet evolution
  • Primary trade-off: Ephemeral runners maximize isolation; warm pools reduce startup time at greater cost and risk.
BEFORE THE NEXT QUESTION

Can you redraw it from memory?

  • Name the source of truth.
  • Trace the write and read paths.
  • Defend one trade-off.
  • Recover from one failure.