04QuestionsBatch inference system

04 · Worked prompt

Batch inference system

Design large job admission, GPU-aware scheduling, checkpointing, artifacts, and cost-aware throughput.
40 minInterview blueprint
INTERVIEW RUBRIC

What a passing answer must show

100 points · 45 minutes

  1. 20pts

    Scope the problem

    0–5 min

    Prioritize the core flows, state the scale, and name the non-goals.

  2. 15pts

    Define contracts

    5–10 min

    Identify durable entities, APIs, idempotency, and the source of truth.

  3. 30pts

    Complete the diagram

    10–25 min

    Trace one write path and one read path. Label the commit boundary and async work.

  4. 20pts

    Lead one deep dive

    25–38 min

    Choose the highest-risk trade-off and explain the mechanism, alternative, and cost.

  5. 15pts

    Prove reliability

    38–45 min

    Walk a failure, recovery, metric, bottleneck, and evolution path.

DRAW THIS FIRST

One complete box-and-arrow design

Design large job admission, GPU-aware scheduling, checkpointing, artifacts, and cost-aware throughput.

Batch inference system · system architecture
Batch inference system system architecture. The job manifest is truth; deterministic shards, checkpoints, and GPU-aware leases make work resumable. Request path: ML users to Batch API to Job orchestrator to Job metadata DB. Asynchronous path: GPU-class queues to GPU scheduler. Read path: Job status API to Progress/output manifest. External dependency: Object/model storage.

Write

ML users enters through Batch API. Job orchestrator owns validation and commits the durable record to Job metadata DB.

Propagate

GPU-class queues separates the committed write from background work. GPU scheduler can retry safely while it builds Progress/output manifest.

Read

Job status API serves from Progress/output manifest, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Object/model storage as an explicit dependency.

Say this first: The job manifest is truth; deterministic shards, checkpoints, and GPU-aware leases make work resumable.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

The job manifest is truth; deterministic shards, checkpoints, and GPU-aware leases make work resumable.

INTERVIEW CONTRACT

Design large job admission, GPU-aware scheduling, checkpointing, artifacts, and cost-aware throughput.

CAPACITY QUESTIONS TO QUANTIFY

Hour jobs · heterogeneous GPUs · throughput focus · tenant budgets · large artifacts. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    ML users → Batch API

    Submit job + inspect output enters over HTTPS / RPC. Batch API handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Batch API → Job orchestrator → Job metadata DB

    Job orchestrator receives the command, checks invariants and retry identity, then uses job + shard plan to update Job metadata DB. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Job orchestrator → GPU-class queues → GPU scheduler → Progress/output manifest

    Job orchestrator emits publish after commit; GPU scheduler uses consume and project / update to build Progress/output manifest. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Batch API → Job status API → Progress/output manifest / Job metadata DB

    Job status API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    GPU scheduler → Object/model storage

    range read/write crosses into Object/model storage. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Batch APIAuth + budget admissionIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Job orchestratorValidate + split deterministic shardsWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Job metadata DBShard state + attemptsAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
GPU-class queuesModel + accelerator + priorityDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
GPU schedulerLease, run, checkpointReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Progress/output manifestCommitted shard artifactsRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Job status APIWeighted progress + diagnosticsRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Object/model storageInputs, checkpoints, outputsExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
PostgreSQL stores jobs/shards/attempts/manifests; S3 stores inputs/checkpoints/outputs; priority queues feed GPU work.
Partitioning / sharding
Partition inputs into immutable ranges/objects; queue by model/hardware class and tenant priority; colocate metadata by job_id.
Indexes
Jobs by tenant/state/time, unique shard identity, lease expiry, and output parts by job/shard/attempt.
Replication + consistency
Job state/accepted outputs use conditional commits. Execution is at least once; output becomes visible only through a complete manifest.
Cache, queue + recovery
Workers checkpoint, heartbeat, retry, and dead-letter; cache weights and hot input blocks locally.
Capacity math
Estimate input TB/job, records/sec/GPU, model GB, GPU memory, job concurrency, completion target, and checkpoint bytes.
Alternative rejected
One huge task minimizes orchestration but loses progress and parallelism; immutable shards plus manifests make retry safe.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

GPU packing

Group model, precision, memory, and sequence-length compatible work while enforcing tenant virtual time

Utilization without fairness starves small jobs
Checkpointing

Checkpoint at deterministic input boundaries and write immutable partial outputs

Opaque process snapshots are hard to move across GPU types
Stragglers

Speculatively duplicate only the tail after measuring expected gain and dedupe output commit

Blind duplication doubles expensive work
05

Failure pressure test

Show detection, containment, recovery, and evidence.

Spot interruption

Resume shard from last committed cursor on a compatible worker

recomputed examples
Bad input shard

Quarantine deterministic data error without retry storm

permanent shard failures
Result commit race

Conditionally publish one output version per shard attempt

discarded late outputs
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution
SAY THIS WHILE YOU DRAW

A four-part talk track

  1. Scope

    “I’ll prioritize submit large datasets and model versions and maximize accelerator throughput fairly.”

  2. Scale

    “The design changes around hour jobs · heterogeneous gpus · throughput focus · tenant budgets · large artifacts.”

  3. Decision

    “Large batches maximize throughput; smaller shards improve fairness and retry cost.”

  4. Risk

    “The first failure I want to pressure-test is: Preemption and stragglers can repeat expensive work or delay the tail.”

Reference details

Open these only after you can explain the diagram above without reading.

01Requirements and state lifecycle4 requirements
  • Submit large datasets and model versions
  • Maximize accelerator throughput fairly
  • Resume after preemption or worker loss
  • Publish versioned outputs with cost and progress
Batch inference system · state lifecycle
02Data model and APIs4 entities · 3 interfaces

Core entities

InferenceJobjob_id, tenant, model_version, input_manifest, stateOwner: Job service
WorkShardjob_id, shard_id, input_range, attempt, stateOwner: Scheduler
Checkpointshard_id, cursor, model_digest, artifact_refOwner: Checkpoint store
OutputManifestjob_id, shard_outputs, schema, versionOwner: Result catalog

External interfaces

POST /v1/inference_jobs

Submit model, input manifest, and output contract

POST /v1/inference_jobs/{id}/cancel

Request durable cancellation

GET /v1/inference_jobs/{id}

Read progress, cost, failures, and result manifest

03Deep dives and trade-offsChoose one

GPU packing

Group model, precision, memory, and sequence-length compatible work while enforcing tenant virtual time

Utilization without fairness starves small jobs

Checkpointing

Checkpoint at deterministic input boundaries and write immutable partial outputs

Opaque process snapshots are hard to move across GPU types

Stragglers

Speculatively duplicate only the tail after measuring expected gain and dedupe output commit

Blind duplication doubles expensive work
04Failures, recovery, and evidence3 scenarios

Spot interruption

Resume shard from last committed cursor on a compatible worker

recomputed examples

Bad input shard

Quarantine deterministic data error without retry storm

permanent shard failures

Result commit race

Conditionally publish one output version per shard attempt

discarded late outputs
05What makes the answer seniorInterviewer signals
  • Throughput, queue time, and cost are separate success metrics
  • Make model and data versions immutable for reproducibility
  • The output manifest is the job’s commit boundary
  • Primary trade-off: Large batches maximize throughput; smaller shards improve fairness and retry cost.
BEFORE THE NEXT QUESTION

Can you redraw it from memory?

  • Name the source of truth.
  • Trace the write and read paths.
  • Defend one trade-off.
  • Recover from one failure.