04QuestionsRemote Devbox
04 · Worked prompt
Remote Devbox
Design secure workspace provisioning, lifecycle control, persistent volumes, terminals, and low-latency streaming.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design secure workspace provisioning, lifecycle control, persistent volumes, terminals, and low-latency streaming.

Write
IDE + terminal enters through Session gateway. Workspace control plane owns validation and commits the durable record to Workspace DB.
Propagate
Provisioning queue separates the committed write from background work. Host agents can retry safely while it builds Fleet/lease index.
Read
Session resolver serves from Fleet/lease index, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Compute + volumes as an explicit dependency.
Say this first: A control-plane record owns desired workspace state; hosts and sessions are leased, fenced runtime resources.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
A control-plane record owns desired workspace state; hosts and sessions are leased, fenced runtime resources.
Design secure workspace provisioning, lifecycle control, persistent volumes, terminals, and low-latency streaming.
Interactive latency · bursty starts · persistent volumes · tenant isolation. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
IDE + terminal → Session gatewayCreate, connect, suspend enters over HTTPS / RPC. Session gateway handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Session gateway → Workspace control plane → Workspace DBWorkspace control plane receives the command, checks invariants and retry identity, then uses commit to update Workspace DB. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Workspace control plane → Provisioning queue → Host agents → Fleet/lease indexWorkspace control plane emits publish after commit; Host agents uses consume and project / update to build Fleet/lease index. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Session gateway → Session resolver → Fleet/lease index / Workspace DBSession resolver uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
Host agents → Compute + volumesfenced attach crosses into Compute + volumes. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
- 06
Deliver without changing the source of truth
Host agents → Terminal/IDE stream → IDE + terminalHost agents uses fan-out; Terminal/IDE stream returns updates over WebSocket / QUIC. Sequence IDs, reconnect cursors, and backpressure make delivery resumable without turning a socket into durable state.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Session gatewayAuth + encrypted proxy | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Workspace control planePlacement + lifecycle state | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Workspace DBDesired state + lease epoch | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| Provisioning queueRegion/resource class | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| Host agentsBoot, attach, heartbeat | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Fleet/lease indexCapacity + live endpoints | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Session resolverWorkspace to active lease | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Compute + volumesVM/container + persistent disk | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
| Terminal/IDE streamLow-latency bidirectional bytes | Connection and delivery state | Separates open connections and fan-out pressure from durable domain state. | Reconnect, ordering, slow consumers |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- PostgreSQL stores workspace lifecycle/leases; object storage stores snapshots; block volumes store active filesystems; Redis routes sessions.
- Partitioning / sharding
- Partition control state by organization/workspace and place compute/volumes by region/capacity pool. One workspace has one fenced owner.
- Indexes
- Unique create key, workspaces by owner/state, leases by host/expiry, and snapshots by workspace/time.
- Replication + consistency
- Lifecycle and fencing epochs are strong. Terminal delivery and capacity inventory may be eventually refreshed; no runner holds unique state.
- Cache, queue + recovery
- Durable stage queues support provision, checkpoint, suspend, resume, and compensating cleanup. Cache image layers near compute.
- Capacity math
- Estimate concurrent workspaces, vCPU/RAM, image GB, snapshot bandwidth, startup p95, idle duration, and regional headroom.
- Alternative rejected
- Orchestrator-only state hides repair; a durable workspace record plus fenced stages makes retries and cleanup inspectable.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
Volume fencing
Increment attachment epoch and reject stale-host writes
Compute failover without storage fencing risks corruptionWarm capacity
Maintain small image-aware pools with strict scrub and tenant reassignment
Fast starts cannot weaken isolationInteractive gateway
Keep gateways stateless beyond connection routing and support reconnect tokens
A gateway failure should not destroy the workspaceFailure pressure test
Show detection, containment, recovery, and evidence.
Host lost
Fence old lease, attach volume elsewhere, and replay agent state
workspace recovery timeAgent heartbeat lost
Distinguish network partition from stopped VM before reprovisioning
false recovery rateIdle leak
Suspend under explicit policy and preserve user notification
idle compute hours- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize create reproducible workspaces quickly and provide low-latency terminal and ide sessions.”
- Scale
“The design changes around interactive latency · bursty starts · persistent volumes · tenant isolation.”
- Decision
“Warm pools reduce startup time but expand cost and reassignment risk.”
- Risk
“The first failure I want to pressure-test is: Zombie compute and volume races can leak data or strand workspaces.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Create reproducible workspaces quickly
- Provide low-latency terminal and IDE sessions
- Persist user files across compute loss
- Enforce tenant isolation and idle-cost controls
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 3 interfaces
Core entities
workspace_id, owner_id, image, desired_state, versionOwner: Control planeworkspace_id, host_id, token, expires_atOwner: Provisionervolume_id, workspace_id, host_id, epochOwner: Storage controlsession_id, workspace_id, gateway, expires_atOwner: Access gatewayExternal interfaces
/v1/workspacesCreate desired state from an image and size
/v1/workspaces/{id}/sessionsMint a short-lived interactive session
/v1/workspaces/{id}Suspend, resume, resize, or delete conditionally
03Deep dives and trade-offsChoose one
Volume fencing
Increment attachment epoch and reject stale-host writes
Compute failover without storage fencing risks corruptionWarm capacity
Maintain small image-aware pools with strict scrub and tenant reassignment
Fast starts cannot weaken isolationInteractive gateway
Keep gateways stateless beyond connection routing and support reconnect tokens
A gateway failure should not destroy the workspace04Failures, recovery, and evidence3 scenarios
Host lost
Fence old lease, attach volume elsewhere, and replay agent state
workspace recovery timeAgent heartbeat lost
Distinguish network partition from stopped VM before reprovisioning
false recovery rateIdle leak
Suspend under explicit policy and preserve user notification
idle compute hours05What makes the answer seniorInterviewer signals
- Draw control plane, data plane, and storage boundary separately
- Persistent identity belongs to the workspace, not the VM
- Explain how stale hosts are prevented from writing after failover
- Primary trade-off: Warm pools reduce startup time but expand cost and reassignment risk.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.