04QuestionsDesign Sora

04 · Worked prompt

Design Sora

Design asynchronous GPU workflows, prompt moderation, progress updates, durable media, and capacity controls.
40 minInterview blueprint
INTERVIEW RUBRIC

What a passing answer must show

100 points · 45 minutes

  1. 20pts

    Scope the problem

    0–5 min

    Prioritize the core flows, state the scale, and name the non-goals.

  2. 15pts

    Define contracts

    5–10 min

    Identify durable entities, APIs, idempotency, and the source of truth.

  3. 30pts

    Complete the diagram

    10–25 min

    Trace one write path and one read path. Label the commit boundary and async work.

  4. 20pts

    Lead one deep dive

    25–38 min

    Choose the highest-risk trade-off and explain the mechanism, alternative, and cost.

  5. 15pts

    Prove reliability

    38–45 min

    Walk a failure, recovery, metric, bottleneck, and evolution path.

DRAW THIS FIRST

One complete box-and-arrow design

Design asynchronous GPU workflows, prompt moderation, progress updates, durable media, and capacity controls.

Design Sora · system architecture
Design Sora system architecture. A durable workflow owns policy and stage state; GPU attempts checkpoint and publish only approved media. Request path: Creation clients to Generation API to Workflow orchestrator to Workflow DB. Asynchronous path: GPU workflow queues to GPU media workers. Read path: Generation status API to Progress/media manifest. External dependency: Object/model storage. Delivery path: Progress stream + CDN.

Write

Creation clients enters through Generation API. Workflow orchestrator owns validation and commits the durable record to Workflow DB.

Propagate

GPU workflow queues separates the committed write from background work. GPU media workers can retry safely while it builds Progress/media manifest.

Read

Generation status API serves from Progress/media manifest, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Object/model storage as an explicit dependency.

Say this first: A durable workflow owns policy and stage state; GPU attempts checkpoint and publish only approved media.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

A durable workflow owns policy and stage state; GPU attempts checkpoint and publish only approved media.

INTERVIEW CONTRACT

Design asynchronous GPU workflows, prompt moderation, progress updates, durable media, and capacity controls.

CAPACITY QUESTIONS TO QUANTIFY

Minute jobs · scarce GPUs · large outputs · queue SLO · user cost budgets. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Creation clients → Generation API

    Prompt, cancel, progress enters over HTTPS / RPC. Generation API handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Generation API → Workflow orchestrator → Workflow DB

    Workflow orchestrator receives the command, checks invariants and retry identity, then uses create workflow to update Workflow DB. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Workflow orchestrator → GPU workflow queues → GPU media workers → Progress/media manifest

    Workflow orchestrator emits publish after commit; GPU media workers uses consume and project / update to build Progress/media manifest. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Generation API → Generation status API → Progress/media manifest / Workflow DB

    Generation status API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    GPU media workers → Object/model storage

    checkpoint / asset crosses into Object/model storage. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

  6. 06
    Deliver without changing the source of truth
    GPU media workers → Progress stream + CDN → Creation clients

    GPU media workers uses fan-out; Progress stream + CDN returns updates over SSE then signed URL. Sequence IDs, reconnect cursors, and backpressure make delivery resumable without turning a socket into durable state.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Generation APIAuth + budget + streamIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Workflow orchestratorModerate + stage state machineWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Workflow DBPolicy, stages, attemptsAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
GPU workflow queuesModel + stage + priorityDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
GPU media workersGenerate, checkpoint, encodeReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Progress/media manifestVersioned approved outputsRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Generation status APIProgress + approved assetsRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Object/model storageCheckpoints, models, mediaExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
Progress stream + CDNEvents then signed assetConnection and delivery stateSeparates open connections and fan-out pressure from durable domain state.Reconnect, ordering, slow consumers
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
PostgreSQL stores jobs/moderation/attempts/manifests; S3/CDN stores inputs/videos; priority queues feed GPU stages.
Partitioning / sharding
Partition metadata by user/job and queues by model/GPU class/priority. Key assets by job/attempt/content digest.
Indexes
Unique client_request_id, jobs by owner/time/state, attempts by lease expiry, and assets by job/version.
Replication + consistency
Lifecycle/billing are strong. Progress is eventual; an asset becomes readable only after checksum, safety, and manifest commit.
Cache, queue + recovery
Durable staged workflows support checkpoint, cancellation fencing, retry, and progress sequence IDs; cache weights on GPU hosts.
Capacity math
Estimate GPU-minutes/job, concurrency, model/checkpoint GB, upload/output bandwidth, queue wait, and cancellation waste.
Alternative rejected
Long-held HTTP requests couple work to connections; durable jobs plus progress streams survive disconnect and capacity waits.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Workflow durability

Persist stage transitions and deterministic seeds so safe stages can resume

A browser connection cannot own a ten-minute GPU job
Capacity

Schedule by GPU memory, model residency, expected duration, and tenant budget

Simple FIFO wastes accelerators and creates unpredictable cost
Moderation boundary

Check prompt before spend and final media before publication; keep rejected assets inaccessible

Safety after CDN publication is too late
05

Failure pressure test

Show detection, containment, recovery, and evidence.

GPU preempted

Resume from compatible checkpoint or restart named stage

recomputed GPU seconds
Cancellation race

Fence later stage commits after cancel version

post-cancel compute and publications
Output blocked

Retain only policy-allowed audit metadata and release capacity

blocked asset exposure count
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution
SAY THIS WHILE YOU DRAW

A four-part talk track

  1. Scope

    “I’ll prioritize submit prompts and generation settings and queue scarce gpu work fairly.”

  2. Scale

    “The design changes around minute jobs · scarce gpus · large outputs · queue slo · user cost budgets.”

  3. Decision

    “Durable asynchronous workflows survive long queues and client disconnects.”

  4. Risk

    “The first failure I want to pressure-test is: GPU loss, unsafe intermediate output, and cancellation races waste capacity or expose content.”

Reference details

Open these only after you can explain the diagram above without reading.

01Requirements and state lifecycle4 requirements
  • Submit prompts and generation settings
  • Queue scarce GPU work fairly
  • Show progress and support cancellation
  • Moderate inputs and outputs and deliver large media safely
Design Sora · state lifecycle
02Data model and APIs4 entities · 3 interfaces

Core entities

GenerationJobjob_id, owner, prompt_ref, model_version, stateOwner: Job service
StageAttemptjob_id, stage, attempt, checkpoint_refOwner: Workflow engine
MediaAssetasset_id, digest, format, moderation_stateOwner: Media service
CapacityReservationtenant, gpu_class, budget, leaseOwner: Scheduler

External interfaces

POST /v1/video_generations

Create a moderated idempotent job

GET /v1/video_generations/{id}/events

Stream queue, stage, preview, and terminal events

POST /v1/video_generations/{id}/cancel

Request cancellation and cost cutoff

03Deep dives and trade-offsChoose one

Workflow durability

Persist stage transitions and deterministic seeds so safe stages can resume

A browser connection cannot own a ten-minute GPU job

Capacity

Schedule by GPU memory, model residency, expected duration, and tenant budget

Simple FIFO wastes accelerators and creates unpredictable cost

Moderation boundary

Check prompt before spend and final media before publication; keep rejected assets inaccessible

Safety after CDN publication is too late
04Failures, recovery, and evidence3 scenarios

GPU preempted

Resume from compatible checkpoint or restart named stage

recomputed GPU seconds

Cancellation race

Fence later stage commits after cancel version

post-cancel compute and publications

Output blocked

Retain only policy-allowed audit metadata and release capacity

blocked asset exposure count
05What makes the answer seniorInterviewer signals
  • The job state machine is part of the product experience
  • Progress should represent durable stages, not invented percentages
  • Discuss cost admission and cancellation—not only GPU scaling
  • Primary trade-off: Durable asynchronous workflows survive long queues and client disconnects.
BEFORE THE NEXT QUESTION

Can you redraw it from memory?

  • Name the source of truth.
  • Trace the write and read paths.
  • Defend one trade-off.
  • Recover from one failure.