04QuestionsBatch inference system
04 · Worked prompt
Batch inference system
Design large job admission, GPU-aware scheduling, checkpointing, artifacts, and cost-aware throughput.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design large job admission, GPU-aware scheduling, checkpointing, artifacts, and cost-aware throughput.

Write
ML users enters through Batch API. Job orchestrator owns validation and commits the durable record to Job metadata DB.
Propagate
GPU-class queues separates the committed write from background work. GPU scheduler can retry safely while it builds Progress/output manifest.
Read
Job status API serves from Progress/output manifest, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Object/model storage as an explicit dependency.
Say this first: The job manifest is truth; deterministic shards, checkpoints, and GPU-aware leases make work resumable.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
The job manifest is truth; deterministic shards, checkpoints, and GPU-aware leases make work resumable.
Design large job admission, GPU-aware scheduling, checkpointing, artifacts, and cost-aware throughput.
Hour jobs · heterogeneous GPUs · throughput focus · tenant budgets · large artifacts. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
ML users → Batch APISubmit job + inspect output enters over HTTPS / RPC. Batch API handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Batch API → Job orchestrator → Job metadata DBJob orchestrator receives the command, checks invariants and retry identity, then uses job + shard plan to update Job metadata DB. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Job orchestrator → GPU-class queues → GPU scheduler → Progress/output manifestJob orchestrator emits publish after commit; GPU scheduler uses consume and project / update to build Progress/output manifest. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Batch API → Job status API → Progress/output manifest / Job metadata DBJob status API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
GPU scheduler → Object/model storagerange read/write crosses into Object/model storage. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Batch APIAuth + budget admission | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Job orchestratorValidate + split deterministic shards | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Job metadata DBShard state + attempts | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| GPU-class queuesModel + accelerator + priority | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| GPU schedulerLease, run, checkpoint | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Progress/output manifestCommitted shard artifacts | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Job status APIWeighted progress + diagnostics | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Object/model storageInputs, checkpoints, outputs | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- PostgreSQL stores jobs/shards/attempts/manifests; S3 stores inputs/checkpoints/outputs; priority queues feed GPU work.
- Partitioning / sharding
- Partition inputs into immutable ranges/objects; queue by model/hardware class and tenant priority; colocate metadata by job_id.
- Indexes
- Jobs by tenant/state/time, unique shard identity, lease expiry, and output parts by job/shard/attempt.
- Replication + consistency
- Job state/accepted outputs use conditional commits. Execution is at least once; output becomes visible only through a complete manifest.
- Cache, queue + recovery
- Workers checkpoint, heartbeat, retry, and dead-letter; cache weights and hot input blocks locally.
- Capacity math
- Estimate input TB/job, records/sec/GPU, model GB, GPU memory, job concurrency, completion target, and checkpoint bytes.
- Alternative rejected
- One huge task minimizes orchestration but loses progress and parallelism; immutable shards plus manifests make retry safe.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
GPU packing
Group model, precision, memory, and sequence-length compatible work while enforcing tenant virtual time
Utilization without fairness starves small jobsCheckpointing
Checkpoint at deterministic input boundaries and write immutable partial outputs
Opaque process snapshots are hard to move across GPU typesStragglers
Speculatively duplicate only the tail after measuring expected gain and dedupe output commit
Blind duplication doubles expensive workFailure pressure test
Show detection, containment, recovery, and evidence.
Spot interruption
Resume shard from last committed cursor on a compatible worker
recomputed examplesBad input shard
Quarantine deterministic data error without retry storm
permanent shard failuresResult commit race
Conditionally publish one output version per shard attempt
discarded late outputs- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize submit large datasets and model versions and maximize accelerator throughput fairly.”
- Scale
“The design changes around hour jobs · heterogeneous gpus · throughput focus · tenant budgets · large artifacts.”
- Decision
“Large batches maximize throughput; smaller shards improve fairness and retry cost.”
- Risk
“The first failure I want to pressure-test is: Preemption and stragglers can repeat expensive work or delay the tail.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Submit large datasets and model versions
- Maximize accelerator throughput fairly
- Resume after preemption or worker loss
- Publish versioned outputs with cost and progress
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 3 interfaces
Core entities
job_id, tenant, model_version, input_manifest, stateOwner: Job servicejob_id, shard_id, input_range, attempt, stateOwner: Schedulershard_id, cursor, model_digest, artifact_refOwner: Checkpoint storejob_id, shard_outputs, schema, versionOwner: Result catalogExternal interfaces
/v1/inference_jobsSubmit model, input manifest, and output contract
/v1/inference_jobs/{id}/cancelRequest durable cancellation
/v1/inference_jobs/{id}Read progress, cost, failures, and result manifest
03Deep dives and trade-offsChoose one
GPU packing
Group model, precision, memory, and sequence-length compatible work while enforcing tenant virtual time
Utilization without fairness starves small jobsCheckpointing
Checkpoint at deterministic input boundaries and write immutable partial outputs
Opaque process snapshots are hard to move across GPU typesStragglers
Speculatively duplicate only the tail after measuring expected gain and dedupe output commit
Blind duplication doubles expensive work04Failures, recovery, and evidence3 scenarios
Spot interruption
Resume shard from last committed cursor on a compatible worker
recomputed examplesBad input shard
Quarantine deterministic data error without retry storm
permanent shard failuresResult commit race
Conditionally publish one output version per shard attempt
discarded late outputs05What makes the answer seniorInterviewer signals
- Throughput, queue time, and cost are separate success metrics
- Make model and data versions immutable for reproducibility
- The output manifest is the job’s commit boundary
- Primary trade-off: Large batches maximize throughput; smaller shards improve fairness and retry cost.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.