04QuestionsChatGPT Playground

04 · Worked prompt

ChatGPT Playground

Design low-latency token streaming, conversation state, model routing, safety checks, and usage metering.
35 minInterview blueprint
INTERVIEW RUBRIC

What a passing answer must show

100 points · 45 minutes

  1. 20pts

    Scope the problem

    0–5 min

    Prioritize the core flows, state the scale, and name the non-goals.

  2. 15pts

    Define contracts

    5–10 min

    Identify durable entities, APIs, idempotency, and the source of truth.

  3. 30pts

    Complete the diagram

    10–25 min

    Trace one write path and one read path. Label the commit boundary and async work.

  4. 20pts

    Lead one deep dive

    25–38 min

    Choose the highest-risk trade-off and explain the mechanism, alternative, and cost.

  5. 15pts

    Prove reliability

    38–45 min

    Walk a failure, recovery, metric, bottleneck, and evolution path.

DRAW THIS FIRST

One complete box-and-arrow design

Design low-latency token streaming, conversation state, model routing, safety checks, and usage metering.

ChatGPT Playground · system architecture
ChatGPT Playground system architecture. Conversation state is durable; generation is scheduled GPU work with a resumable token stream. Request path: Playground UI to API + stream edge to Generation service to Conversation DB. Asynchronous path: Inference queue to GPU scheduler. Read path: Conversation API to Generation cache. External dependency: Model serving fleet. Delivery path: Token stream.

Write

Playground UI enters through API + stream edge. Generation service owns validation and commits the durable record to Conversation DB.

Propagate

Inference queue separates the committed write from background work. GPU scheduler can retry safely while it builds Generation cache.

Read

Conversation API serves from Generation cache, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Model serving fleet as an explicit dependency.

Say this first: Conversation state is durable; generation is scheduled GPU work with a resumable token stream.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

Conversation state is durable; generation is scheduled GPU work with a resumable token stream.

INTERVIEW CONTRACT

Design low-latency token streaming, conversation state, model routing, safety checks, and usage metering.

CAPACITY QUESTIONS TO QUANTIFY

High concurrency · sub-second TTFT · variable prompts · per-tenant quotas. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Playground UI → API + stream edge

    Prompt + stream cursor enters over HTTPS / RPC. API + stream edge handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    API + stream edge → Generation service → Conversation DB

    Generation service receives the POST generation, checks invariants and retry identity, then uses transaction to update Conversation DB. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Generation service → Inference queue → GPU scheduler → Generation cache

    Generation service emits publish after commit; GPU scheduler uses consume and project / update to build Generation cache. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    API + stream edge → Conversation API → Generation cache / Conversation DB

    Conversation API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    GPU scheduler → Model serving fleet

    dispatch crosses into Model serving fleet. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

  6. 06
    Deliver without changing the source of truth
    GPU scheduler → Token stream → Playground UI

    GPU scheduler uses token deltas; Token stream returns updates over SSE / WS. Sequence IDs, reconnect cursors, and backpressure make delivery resumable without turning a socket into durable state.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
API + stream edgeAuth, quota, SSEIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Generation servicePersist + safety + routeWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Conversation DBMessages and usageAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Inference queueModel + priority classDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
GPU schedulerBatch, prefill, decodeReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Generation cachePartial output + cursorRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Conversation APIMessages + generation stateRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Model serving fleetGPU workers + KV cacheExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
Token streamOrdered deltas + terminal eventConnection and delivery stateSeparates open connections and fan-out pressure from durable domain state.Reconnect, ordering, slow consumers
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
PostgreSQL stores conversations, messages, generation state, and the append-only usage ledger; object storage holds attachments; Redis holds ephemeral stream/admission state.
Partitioning / sharding
Start unsharded; later colocate by account_id or conversation_id so ordering and authorization stay local. Partition usage by account and time.
Indexes
Use (conversation_id, created_at, message_id), unique request/idempotency keys, and (account_id, created_at) for usage scans.
Replication + consistency
Messages and usage commit synchronously across zones. Token events are ordered per generation and resumable; partial output may lag the durable cursor.
Cache, queue + recovery
Redis stores quota counters and short replay buffers. A priority queue partitions by model/capacity class; GPU workers are leased and replaceable.
Capacity math
Estimate concurrent streams, prompt/output tokens/sec, KV-cache bytes/token, TTFT, inter-token latency, and GPU utilization.
Alternative rejected
Whole-conversation documents are convenient but become contended and expensive to rewrite; relational message rows expose ordering and idempotency.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Streaming contract

Use SSE with ordered event IDs and replay from the last acknowledged cursor

A reconnect should not start a second billable generation
Scheduling

Use continuous batching with prompt-length-aware admission and tenant fairness

Throughput cannot be optimized at the expense of interactive tail latency
Message commit

Store partial output as a generation artifact, then promote it to a completed message only at a terminal boundary

Conversation truth and transport state should not be conflated
05

Failure pressure test

Show detection, containment, recovery, and evidence.

Client disconnect

Continue briefly and allow cursor resume or cancel by policy

orphan generation rate and resumed streams
GPU failure

Retry only before externally visible nondeterministic output or create a new generation version

failed token position and duplicate usage
Safety service timeout

Fail closed for creation and expose a recoverable state

safety dependency p99 and blocked requests
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution
SAY THIS WHILE YOU DRAW

A four-part talk track

  1. Scope

    “I’ll prioritize create and resume conversations and stream tokens with low time to first token.”

  2. Scale

    “The design changes around high concurrency · sub-second ttft · variable prompts · per-tenant quotas.”

  3. Decision

    “SSE is simple for one-way tokens; WebSockets earn bidirectional control.”

  4. Risk

    “The first failure I want to pressure-test is: Partial streams and disconnects must not corrupt message state or usage totals.”

Reference details

Open these only after you can explain the diagram above without reading.

01Requirements and state lifecycle4 requirements
  • Create and resume conversations
  • Stream tokens with low time to first token
  • Support model parameters and saved presets
  • Enforce safety, quotas, and usage accounting
ChatGPT Playground · state lifecycle
02Data model and APIs4 entities · 3 interfaces

Core entities

Conversationconversation_id, owner_id, title, versionOwner: Conversation service
Messagemessage_id, role, content, status, token_countOwner: Conversation service
Generationgeneration_id, model, parameters, stateOwner: Inference control plane
UsageEntryrequest_id, prompt_tokens, output_tokens, costOwner: Usage ledger

External interfaces

POST /v1/conversations/{id}/generations

Create an idempotent generation from a message version

GET /v1/generations/{id}/events

Stream token, safety, usage, and terminal events

POST /v1/generations/{id}/cancel

Request cancellation without deleting partial history

03Deep dives and trade-offsChoose one

Streaming contract

Use SSE with ordered event IDs and replay from the last acknowledged cursor

A reconnect should not start a second billable generation

Scheduling

Use continuous batching with prompt-length-aware admission and tenant fairness

Throughput cannot be optimized at the expense of interactive tail latency

Message commit

Store partial output as a generation artifact, then promote it to a completed message only at a terminal boundary

Conversation truth and transport state should not be conflated
04Failures, recovery, and evidence3 scenarios

Client disconnect

Continue briefly and allow cursor resume or cancel by policy

orphan generation rate and resumed streams

GPU failure

Retry only before externally visible nondeterministic output or create a new generation version

failed token position and duplicate usage

Safety service timeout

Fail closed for creation and expose a recoverable state

safety dependency p99 and blocked requests
05What makes the answer seniorInterviewer signals
  • Name the exact moment the assistant message becomes durable
  • Separate conversation state, inference execution, and usage accounting
  • Discuss TTFT and inter-token latency instead of one generic latency number
  • Primary trade-off: SSE is simple for one-way tokens; WebSockets earn bidirectional control.
BEFORE THE NEXT QUESTION

Can you redraw it from memory?

  • Name the source of truth.
  • Trace the write and read paths.
  • Defend one trade-off.
  • Recover from one failure.