04QuestionsRemote Devbox

04 · Worked prompt

Remote Devbox

Design secure workspace provisioning, lifecycle control, persistent volumes, terminals, and low-latency streaming.
40 minInterview blueprint
INTERVIEW RUBRIC

What a passing answer must show

100 points · 45 minutes

  1. 20pts

    Scope the problem

    0–5 min

    Prioritize the core flows, state the scale, and name the non-goals.

  2. 15pts

    Define contracts

    5–10 min

    Identify durable entities, APIs, idempotency, and the source of truth.

  3. 30pts

    Complete the diagram

    10–25 min

    Trace one write path and one read path. Label the commit boundary and async work.

  4. 20pts

    Lead one deep dive

    25–38 min

    Choose the highest-risk trade-off and explain the mechanism, alternative, and cost.

  5. 15pts

    Prove reliability

    38–45 min

    Walk a failure, recovery, metric, bottleneck, and evolution path.

DRAW THIS FIRST

One complete box-and-arrow design

Design secure workspace provisioning, lifecycle control, persistent volumes, terminals, and low-latency streaming.

Remote Devbox · system architecture
Remote Devbox system architecture. A control-plane record owns desired workspace state; hosts and sessions are leased, fenced runtime resources. Request path: IDE + terminal to Session gateway to Workspace control plane to Workspace DB. Asynchronous path: Provisioning queue to Host agents. Read path: Session resolver to Fleet/lease index. External dependency: Compute + volumes. Delivery path: Terminal/IDE stream.

Write

IDE + terminal enters through Session gateway. Workspace control plane owns validation and commits the durable record to Workspace DB.

Propagate

Provisioning queue separates the committed write from background work. Host agents can retry safely while it builds Fleet/lease index.

Read

Session resolver serves from Fleet/lease index, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Compute + volumes as an explicit dependency.

Say this first: A control-plane record owns desired workspace state; hosts and sessions are leased, fenced runtime resources.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

A control-plane record owns desired workspace state; hosts and sessions are leased, fenced runtime resources.

INTERVIEW CONTRACT

Design secure workspace provisioning, lifecycle control, persistent volumes, terminals, and low-latency streaming.

CAPACITY QUESTIONS TO QUANTIFY

Interactive latency · bursty starts · persistent volumes · tenant isolation. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    IDE + terminal → Session gateway

    Create, connect, suspend enters over HTTPS / RPC. Session gateway handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Session gateway → Workspace control plane → Workspace DB

    Workspace control plane receives the command, checks invariants and retry identity, then uses commit to update Workspace DB. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Workspace control plane → Provisioning queue → Host agents → Fleet/lease index

    Workspace control plane emits publish after commit; Host agents uses consume and project / update to build Fleet/lease index. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Session gateway → Session resolver → Fleet/lease index / Workspace DB

    Session resolver uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    Host agents → Compute + volumes

    fenced attach crosses into Compute + volumes. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

  6. 06
    Deliver without changing the source of truth
    Host agents → Terminal/IDE stream → IDE + terminal

    Host agents uses fan-out; Terminal/IDE stream returns updates over WebSocket / QUIC. Sequence IDs, reconnect cursors, and backpressure make delivery resumable without turning a socket into durable state.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Session gatewayAuth + encrypted proxyIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Workspace control planePlacement + lifecycle stateWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Workspace DBDesired state + lease epochAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Provisioning queueRegion/resource classDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Host agentsBoot, attach, heartbeatReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Fleet/lease indexCapacity + live endpointsRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Session resolverWorkspace to active leaseRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Compute + volumesVM/container + persistent diskExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
Terminal/IDE streamLow-latency bidirectional bytesConnection and delivery stateSeparates open connections and fan-out pressure from durable domain state.Reconnect, ordering, slow consumers
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
PostgreSQL stores workspace lifecycle/leases; object storage stores snapshots; block volumes store active filesystems; Redis routes sessions.
Partitioning / sharding
Partition control state by organization/workspace and place compute/volumes by region/capacity pool. One workspace has one fenced owner.
Indexes
Unique create key, workspaces by owner/state, leases by host/expiry, and snapshots by workspace/time.
Replication + consistency
Lifecycle and fencing epochs are strong. Terminal delivery and capacity inventory may be eventually refreshed; no runner holds unique state.
Cache, queue + recovery
Durable stage queues support provision, checkpoint, suspend, resume, and compensating cleanup. Cache image layers near compute.
Capacity math
Estimate concurrent workspaces, vCPU/RAM, image GB, snapshot bandwidth, startup p95, idle duration, and regional headroom.
Alternative rejected
Orchestrator-only state hides repair; a durable workspace record plus fenced stages makes retries and cleanup inspectable.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Volume fencing

Increment attachment epoch and reject stale-host writes

Compute failover without storage fencing risks corruption
Warm capacity

Maintain small image-aware pools with strict scrub and tenant reassignment

Fast starts cannot weaken isolation
Interactive gateway

Keep gateways stateless beyond connection routing and support reconnect tokens

A gateway failure should not destroy the workspace
05

Failure pressure test

Show detection, containment, recovery, and evidence.

Host lost

Fence old lease, attach volume elsewhere, and replay agent state

workspace recovery time
Agent heartbeat lost

Distinguish network partition from stopped VM before reprovisioning

false recovery rate
Idle leak

Suspend under explicit policy and preserve user notification

idle compute hours
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution
SAY THIS WHILE YOU DRAW

A four-part talk track

  1. Scope

    “I’ll prioritize create reproducible workspaces quickly and provide low-latency terminal and ide sessions.”

  2. Scale

    “The design changes around interactive latency · bursty starts · persistent volumes · tenant isolation.”

  3. Decision

    “Warm pools reduce startup time but expand cost and reassignment risk.”

  4. Risk

    “The first failure I want to pressure-test is: Zombie compute and volume races can leak data or strand workspaces.”

Reference details

Open these only after you can explain the diagram above without reading.

01Requirements and state lifecycle4 requirements
  • Create reproducible workspaces quickly
  • Provide low-latency terminal and IDE sessions
  • Persist user files across compute loss
  • Enforce tenant isolation and idle-cost controls
Remote Devbox · state lifecycle
02Data model and APIs4 entities · 3 interfaces

Core entities

Workspaceworkspace_id, owner_id, image, desired_state, versionOwner: Control plane
InstanceLeaseworkspace_id, host_id, token, expires_atOwner: Provisioner
VolumeAttachmentvolume_id, workspace_id, host_id, epochOwner: Storage control
Sessionsession_id, workspace_id, gateway, expires_atOwner: Access gateway

External interfaces

POST /v1/workspaces

Create desired state from an image and size

POST /v1/workspaces/{id}/sessions

Mint a short-lived interactive session

PATCH /v1/workspaces/{id}

Suspend, resume, resize, or delete conditionally

03Deep dives and trade-offsChoose one

Volume fencing

Increment attachment epoch and reject stale-host writes

Compute failover without storage fencing risks corruption

Warm capacity

Maintain small image-aware pools with strict scrub and tenant reassignment

Fast starts cannot weaken isolation

Interactive gateway

Keep gateways stateless beyond connection routing and support reconnect tokens

A gateway failure should not destroy the workspace
04Failures, recovery, and evidence3 scenarios

Host lost

Fence old lease, attach volume elsewhere, and replay agent state

workspace recovery time

Agent heartbeat lost

Distinguish network partition from stopped VM before reprovisioning

false recovery rate

Idle leak

Suspend under explicit policy and preserve user notification

idle compute hours
05What makes the answer seniorInterviewer signals
  • Draw control plane, data plane, and storage boundary separately
  • Persistent identity belongs to the workspace, not the VM
  • Explain how stale hosts are prevented from writing after failover
  • Primary trade-off: Warm pools reduce startup time but expand cost and reassignment risk.
BEFORE THE NEXT QUESTION

Can you redraw it from memory?

  • Name the source of truth.
  • Trace the write and read paths.
  • Defend one trade-off.
  • Recover from one failure.