04QuestionsCI/CD system
04 · Worked prompt
CI/CD system
Design workflow DAGs, isolated runners, artifacts, logs, retries, and fair scheduling for bursty builds.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design workflow DAGs, isolated runners, artifacts, logs, retries, and fair scheduling for bursty builds.

Write
Developer + Git enters through Webhook/API edge. Workflow orchestrator owns validation and commits the durable record to Workflow DB.
Propagate
Priority task queues separates the committed write from background work. Runner manager can retry safely while it builds Log/artifact index.
Read
Run status API serves from Log/artifact index, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Sandbox runner pool as an explicit dependency.
Say this first: The workflow database owns DAG state; runners are leased, isolated, and disposable.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
The workflow database owns DAG state; runners are leased, isolated, and disposable.
Design workflow DAGs, isolated runners, artifacts, logs, retries, and fair scheduling for bursty builds.
Bursty pushes · thousands concurrent jobs · minute-to-hour runtimes · untrusted code. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
Developer + Git → Webhook/API edgeCommit, webhook, cancel enters over HTTPS / RPC. Webhook/API edge handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Webhook/API edge → Workflow orchestrator → Workflow DBWorkflow orchestrator receives the verified webhook, checks invariants and retry identity, then uses state transition to update Workflow DB. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Workflow orchestrator → Priority task queues → Runner manager → Log/artifact indexWorkflow orchestrator emits publish after commit; Runner manager uses consume and project / update to build Log/artifact index. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Webhook/API edge → Run status API → Log/artifact index / Workflow DBRun status API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
Runner manager → Sandbox runner poollease runner crosses into Sandbox runner pool. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
- 06
Deliver without changing the source of truth
Runner manager → Live log gateway → Developer + GitRunner manager uses fan-out; Live log gateway returns updates over SSE logs. Sequence IDs, reconnect cursors, and backpressure make delivery resumable without turning a socket into durable state.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Webhook/API edgeVerify, auth, quota | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Workflow orchestratorExpand DAG + release tasks | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Workflow DBRun + task state | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| Priority task queuesResource-class partitions | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| Runner managerFenced leases + retries | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Log/artifact indexCursors + signed links | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Run status APIDAG, attempts, artifacts | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Sandbox runner poolCompile, test, build | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
| Live log gatewayTail task output | Connection and delivery state | Separates open connections and fan-out pressure from durable domain state. | Reconnect, ordering, slow consumers |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- PostgreSQL owns workflow/DAG and attempt state, S3 owns logs/artifacts, and Kafka/SQS or a durable priority queue owns runnable work.
- Partitioning / sharding
- Partition metadata by organization_id/repository_id and queue by resource class. Keep one run's DAG local; key objects by org/run/task/digest.
- Indexes
- Unique provider_delivery_id; (run_id, state), (task_id, attempt), (queue_class, priority, ready_at), and artifact digest indexes.
- Replication + consistency
- DAG transitions and accepted terminal attempts use transactions or conditional writes. Logs and artifact indexes may converge asynchronously.
- Cache, queue + recovery
- Use an outbox to avoid DB/queue dual writes; runners use fenced leases, heartbeats, bounded retries, and dead letters.
- Capacity math
- Estimate webhook bursts, concurrent tasks, log MB/sec, artifact GB/day, runner cold start, and queue age per resource class.
- Alternative rejected
- Keeping the DAG only in a broker makes repair opaque; durable relational workflow state plus queues for ready work is recoverable.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
DAG scheduling
Maintain remaining-dependency counts and atomically release a node once
Scanning the whole graph after every task is wasteful and race-proneRunner isolation
Use ephemeral microVMs, outbound network policy, short-lived credentials, and immutable base images
The user workload is hostile by defaultRetry semantics
Create a new attempt under the same logical task and publish artifacts only after terminal commit
Retries must not overwrite evidence from earlier attemptsFailure pressure test
Show detection, containment, recovery, and evidence.
Runner heartbeat lost
Expire the lease and enqueue a new attempt after a fencing check
expired lease count and duplicate runner activityWebhook replay
Deduplicate provider delivery ID and repo commit tuple
webhook dedupe rateLog backpressure
Spool locally with bounded buffers and preserve job execution
dropped log bytes and stream lag- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize trigger workflows from signed repository events and execute dag stages with retries and cancellation.”
- Scale
“The design changes around bursty pushes · thousands concurrent jobs · minute-to-hour runtimes · untrusted code.”
- Decision
“Ephemeral runners maximize isolation; warm pools reduce startup time at greater cost and risk.”
- Risk
“The first failure I want to pressure-test is: Duplicate webhooks and lost heartbeats can execute deployments twice or leak capacity.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Trigger workflows from signed repository events
- Execute DAG stages with retries and cancellation
- Stream logs and preserve artifacts
- Isolate tenants and support bursty workloads
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 3 interfaces
Core entities
repo_id, revision, DAG, policyOwner: Workflow servicerun_id, commit_sha, state, versionOwner: Orchestratortask_id, attempt, lease, resultOwner: Schedulerartifact_id, digest, producer, retentionOwner: Artifact serviceExternal interfaces
/v1/hooks/gitVerify and deduplicate a repository event
/v1/runs/{id}/cancelVersion-check and cancel runnable work
/v1/runs/{id}/eventsStream logs and state changes from a cursor
03Deep dives and trade-offsChoose one
DAG scheduling
Maintain remaining-dependency counts and atomically release a node once
Scanning the whole graph after every task is wasteful and race-proneRunner isolation
Use ephemeral microVMs, outbound network policy, short-lived credentials, and immutable base images
The user workload is hostile by defaultRetry semantics
Create a new attempt under the same logical task and publish artifacts only after terminal commit
Retries must not overwrite evidence from earlier attempts04Failures, recovery, and evidence3 scenarios
Runner heartbeat lost
Expire the lease and enqueue a new attempt after a fencing check
expired lease count and duplicate runner activityWebhook replay
Deduplicate provider delivery ID and repo commit tuple
webhook dedupe rateLog backpressure
Spool locally with bounded buffers and preserve job execution
dropped log bytes and stream lag05What makes the answer seniorInterviewer signals
- Draw the control plane separately from runner data paths
- The hardest invariant is one terminal task result despite at-least-once execution
- Staff-level answers discuss fairness, secrets, and fleet evolution
- Primary trade-off: Ephemeral runners maximize isolation; warm pools reduce startup time at greater cost and risk.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.