04QuestionsChatGPT Playground
04 · Worked prompt
ChatGPT Playground
Design low-latency token streaming, conversation state, model routing, safety checks, and usage metering.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design low-latency token streaming, conversation state, model routing, safety checks, and usage metering.

Write
Playground UI enters through API + stream edge. Generation service owns validation and commits the durable record to Conversation DB.
Propagate
Inference queue separates the committed write from background work. GPU scheduler can retry safely while it builds Generation cache.
Read
Conversation API serves from Generation cache, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Model serving fleet as an explicit dependency.
Say this first: Conversation state is durable; generation is scheduled GPU work with a resumable token stream.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
Conversation state is durable; generation is scheduled GPU work with a resumable token stream.
Design low-latency token streaming, conversation state, model routing, safety checks, and usage metering.
High concurrency · sub-second TTFT · variable prompts · per-tenant quotas. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
Playground UI → API + stream edgePrompt + stream cursor enters over HTTPS / RPC. API + stream edge handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
API + stream edge → Generation service → Conversation DBGeneration service receives the POST generation, checks invariants and retry identity, then uses transaction to update Conversation DB. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Generation service → Inference queue → GPU scheduler → Generation cacheGeneration service emits publish after commit; GPU scheduler uses consume and project / update to build Generation cache. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
API + stream edge → Conversation API → Generation cache / Conversation DBConversation API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
GPU scheduler → Model serving fleetdispatch crosses into Model serving fleet. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
- 06
Deliver without changing the source of truth
GPU scheduler → Token stream → Playground UIGPU scheduler uses token deltas; Token stream returns updates over SSE / WS. Sequence IDs, reconnect cursors, and backpressure make delivery resumable without turning a socket into durable state.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| API + stream edgeAuth, quota, SSE | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Generation servicePersist + safety + route | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Conversation DBMessages and usage | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| Inference queueModel + priority class | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| GPU schedulerBatch, prefill, decode | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Generation cachePartial output + cursor | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Conversation APIMessages + generation state | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Model serving fleetGPU workers + KV cache | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
| Token streamOrdered deltas + terminal event | Connection and delivery state | Separates open connections and fan-out pressure from durable domain state. | Reconnect, ordering, slow consumers |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- PostgreSQL stores conversations, messages, generation state, and the append-only usage ledger; object storage holds attachments; Redis holds ephemeral stream/admission state.
- Partitioning / sharding
- Start unsharded; later colocate by account_id or conversation_id so ordering and authorization stay local. Partition usage by account and time.
- Indexes
- Use (conversation_id, created_at, message_id), unique request/idempotency keys, and (account_id, created_at) for usage scans.
- Replication + consistency
- Messages and usage commit synchronously across zones. Token events are ordered per generation and resumable; partial output may lag the durable cursor.
- Cache, queue + recovery
- Redis stores quota counters and short replay buffers. A priority queue partitions by model/capacity class; GPU workers are leased and replaceable.
- Capacity math
- Estimate concurrent streams, prompt/output tokens/sec, KV-cache bytes/token, TTFT, inter-token latency, and GPU utilization.
- Alternative rejected
- Whole-conversation documents are convenient but become contended and expensive to rewrite; relational message rows expose ordering and idempotency.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
Streaming contract
Use SSE with ordered event IDs and replay from the last acknowledged cursor
A reconnect should not start a second billable generationScheduling
Use continuous batching with prompt-length-aware admission and tenant fairness
Throughput cannot be optimized at the expense of interactive tail latencyMessage commit
Store partial output as a generation artifact, then promote it to a completed message only at a terminal boundary
Conversation truth and transport state should not be conflatedFailure pressure test
Show detection, containment, recovery, and evidence.
Client disconnect
Continue briefly and allow cursor resume or cancel by policy
orphan generation rate and resumed streamsGPU failure
Retry only before externally visible nondeterministic output or create a new generation version
failed token position and duplicate usageSafety service timeout
Fail closed for creation and expose a recoverable state
safety dependency p99 and blocked requests- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize create and resume conversations and stream tokens with low time to first token.”
- Scale
“The design changes around high concurrency · sub-second ttft · variable prompts · per-tenant quotas.”
- Decision
“SSE is simple for one-way tokens; WebSockets earn bidirectional control.”
- Risk
“The first failure I want to pressure-test is: Partial streams and disconnects must not corrupt message state or usage totals.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Create and resume conversations
- Stream tokens with low time to first token
- Support model parameters and saved presets
- Enforce safety, quotas, and usage accounting
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 3 interfaces
Core entities
conversation_id, owner_id, title, versionOwner: Conversation servicemessage_id, role, content, status, token_countOwner: Conversation servicegeneration_id, model, parameters, stateOwner: Inference control planerequest_id, prompt_tokens, output_tokens, costOwner: Usage ledgerExternal interfaces
/v1/conversations/{id}/generationsCreate an idempotent generation from a message version
/v1/generations/{id}/eventsStream token, safety, usage, and terminal events
/v1/generations/{id}/cancelRequest cancellation without deleting partial history
03Deep dives and trade-offsChoose one
Streaming contract
Use SSE with ordered event IDs and replay from the last acknowledged cursor
A reconnect should not start a second billable generationScheduling
Use continuous batching with prompt-length-aware admission and tenant fairness
Throughput cannot be optimized at the expense of interactive tail latencyMessage commit
Store partial output as a generation artifact, then promote it to a completed message only at a terminal boundary
Conversation truth and transport state should not be conflated04Failures, recovery, and evidence3 scenarios
Client disconnect
Continue briefly and allow cursor resume or cancel by policy
orphan generation rate and resumed streamsGPU failure
Retry only before externally visible nondeterministic output or create a new generation version
failed token position and duplicate usageSafety service timeout
Fail closed for creation and expose a recoverable state
safety dependency p99 and blocked requests05What makes the answer seniorInterviewer signals
- Name the exact moment the assistant message becomes durable
- Separate conversation state, inference execution, and usage accounting
- Discuss TTFT and inter-token latency instead of one generic latency number
- Primary trade-off: SSE is simple for one-way tokens; WebSockets earn bidirectional control.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.