04QuestionsDesign Sora
04 · Worked prompt
Design Sora
Design asynchronous GPU workflows, prompt moderation, progress updates, durable media, and capacity controls.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design asynchronous GPU workflows, prompt moderation, progress updates, durable media, and capacity controls.

Write
Creation clients enters through Generation API. Workflow orchestrator owns validation and commits the durable record to Workflow DB.
Propagate
GPU workflow queues separates the committed write from background work. GPU media workers can retry safely while it builds Progress/media manifest.
Read
Generation status API serves from Progress/media manifest, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Object/model storage as an explicit dependency.
Say this first: A durable workflow owns policy and stage state; GPU attempts checkpoint and publish only approved media.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
A durable workflow owns policy and stage state; GPU attempts checkpoint and publish only approved media.
Design asynchronous GPU workflows, prompt moderation, progress updates, durable media, and capacity controls.
Minute jobs · scarce GPUs · large outputs · queue SLO · user cost budgets. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
Creation clients → Generation APIPrompt, cancel, progress enters over HTTPS / RPC. Generation API handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Generation API → Workflow orchestrator → Workflow DBWorkflow orchestrator receives the command, checks invariants and retry identity, then uses create workflow to update Workflow DB. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Workflow orchestrator → GPU workflow queues → GPU media workers → Progress/media manifestWorkflow orchestrator emits publish after commit; GPU media workers uses consume and project / update to build Progress/media manifest. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Generation API → Generation status API → Progress/media manifest / Workflow DBGeneration status API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
GPU media workers → Object/model storagecheckpoint / asset crosses into Object/model storage. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
- 06
Deliver without changing the source of truth
GPU media workers → Progress stream + CDN → Creation clientsGPU media workers uses fan-out; Progress stream + CDN returns updates over SSE then signed URL. Sequence IDs, reconnect cursors, and backpressure make delivery resumable without turning a socket into durable state.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Generation APIAuth + budget + stream | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Workflow orchestratorModerate + stage state machine | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Workflow DBPolicy, stages, attempts | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| GPU workflow queuesModel + stage + priority | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| GPU media workersGenerate, checkpoint, encode | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Progress/media manifestVersioned approved outputs | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Generation status APIProgress + approved assets | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Object/model storageCheckpoints, models, media | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
| Progress stream + CDNEvents then signed asset | Connection and delivery state | Separates open connections and fan-out pressure from durable domain state. | Reconnect, ordering, slow consumers |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- PostgreSQL stores jobs/moderation/attempts/manifests; S3/CDN stores inputs/videos; priority queues feed GPU stages.
- Partitioning / sharding
- Partition metadata by user/job and queues by model/GPU class/priority. Key assets by job/attempt/content digest.
- Indexes
- Unique client_request_id, jobs by owner/time/state, attempts by lease expiry, and assets by job/version.
- Replication + consistency
- Lifecycle/billing are strong. Progress is eventual; an asset becomes readable only after checksum, safety, and manifest commit.
- Cache, queue + recovery
- Durable staged workflows support checkpoint, cancellation fencing, retry, and progress sequence IDs; cache weights on GPU hosts.
- Capacity math
- Estimate GPU-minutes/job, concurrency, model/checkpoint GB, upload/output bandwidth, queue wait, and cancellation waste.
- Alternative rejected
- Long-held HTTP requests couple work to connections; durable jobs plus progress streams survive disconnect and capacity waits.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
Workflow durability
Persist stage transitions and deterministic seeds so safe stages can resume
A browser connection cannot own a ten-minute GPU jobCapacity
Schedule by GPU memory, model residency, expected duration, and tenant budget
Simple FIFO wastes accelerators and creates unpredictable costModeration boundary
Check prompt before spend and final media before publication; keep rejected assets inaccessible
Safety after CDN publication is too lateFailure pressure test
Show detection, containment, recovery, and evidence.
GPU preempted
Resume from compatible checkpoint or restart named stage
recomputed GPU secondsCancellation race
Fence later stage commits after cancel version
post-cancel compute and publicationsOutput blocked
Retain only policy-allowed audit metadata and release capacity
blocked asset exposure count- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize submit prompts and generation settings and queue scarce gpu work fairly.”
- Scale
“The design changes around minute jobs · scarce gpus · large outputs · queue slo · user cost budgets.”
- Decision
“Durable asynchronous workflows survive long queues and client disconnects.”
- Risk
“The first failure I want to pressure-test is: GPU loss, unsafe intermediate output, and cancellation races waste capacity or expose content.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Submit prompts and generation settings
- Queue scarce GPU work fairly
- Show progress and support cancellation
- Moderate inputs and outputs and deliver large media safely
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 3 interfaces
Core entities
job_id, owner, prompt_ref, model_version, stateOwner: Job servicejob_id, stage, attempt, checkpoint_refOwner: Workflow engineasset_id, digest, format, moderation_stateOwner: Media servicetenant, gpu_class, budget, leaseOwner: SchedulerExternal interfaces
/v1/video_generationsCreate a moderated idempotent job
/v1/video_generations/{id}/eventsStream queue, stage, preview, and terminal events
/v1/video_generations/{id}/cancelRequest cancellation and cost cutoff
03Deep dives and trade-offsChoose one
Workflow durability
Persist stage transitions and deterministic seeds so safe stages can resume
A browser connection cannot own a ten-minute GPU jobCapacity
Schedule by GPU memory, model residency, expected duration, and tenant budget
Simple FIFO wastes accelerators and creates unpredictable costModeration boundary
Check prompt before spend and final media before publication; keep rejected assets inaccessible
Safety after CDN publication is too late04Failures, recovery, and evidence3 scenarios
GPU preempted
Resume from compatible checkpoint or restart named stage
recomputed GPU secondsCancellation race
Fence later stage commits after cancel version
post-cancel compute and publicationsOutput blocked
Retain only policy-allowed audit metadata and release capacity
blocked asset exposure count05What makes the answer seniorInterviewer signals
- The job state machine is part of the product experience
- Progress should represent durable stages, not invented percentages
- Discuss cost admission and cancellation—not only GPU scaling
- Primary trade-off: Durable asynchronous workflows survive long queues and client disconnects.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.