03Building blocksBlob storage

03 · Building block

Blob storage

Separate immutable large objects from metadata and design multipart upload, durability, and lifecycle policies.
8 minConcept guideReference-informed · independently authored
01

Architecture map

See where the component sits in a real system.

Blob storage · system architecture
Blob storage system architecture. Metadata and bytes have separate owners; clients transfer chunks directly using scoped, resumable sessions. Request path: Upload/download clients to Metadata API to Upload coordinator to Metadata database. Asynchronous path: Object event queue to Scan/lifecycle workers. Read path: Download resolver to Object/CDN index. External dependency: Object storage fleet.

Write

Upload/download clients enters through Metadata API. Upload coordinator owns validation and commits the durable record to Metadata database.

Propagate

Object event queue separates the committed write from background work. Scan/lifecycle workers can retry safely while it builds Object/CDN index.

Read

Download resolver serves from Object/CDN index, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Object storage fleet as an explicit dependency.

Say this first: Metadata and bytes have separate owners; clients transfer chunks directly using scoped, resumable sessions.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

Metadata and bytes have separate owners; clients transfer chunks directly using scoped, resumable sessions.

INTERVIEW CONTRACT

Separate immutable large objects from metadata and design multipart upload, durability, and lifecycle policies.

CAPACITY QUESTIONS TO QUANTIFY

Object size distribution · concurrency · part size · retention · durability · egress. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Upload/download clients → Metadata API

    Multipart bytes + ranges enters over HTTPS / RPC. Metadata API handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Metadata API → Upload coordinator → Metadata database

    Upload coordinator receives the command, checks invariants and retry identity, then uses commit manifest to update Metadata database. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Upload coordinator → Object event queue → Scan/lifecycle workers → Object/CDN index

    Upload coordinator emits publish after commit; Scan/lifecycle workers uses consume and project / update to build Object/CDN index. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Metadata API → Download resolver → Object/CDN index / Metadata database

    Download resolver uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    Upload/download clients → Object storage fleet

    direct multipart I/O crosses into Object storage fleet. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Metadata APIAuth + signed sessionsIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Upload coordinatorParts, checksums, finalizeWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Metadata databaseObject version + manifestAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Object event queueFinalize, scan, lifecycleDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Scan/lifecycle workersVerify, tier, expireReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Object/CDN indexLocation + cache metadataRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Download resolverManifest + signed URLRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Object storage fleetErasure-coded immutable chunksExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
S3/GCS-style object storage holds immutable bytes; SQL/distributed SQL holds metadata, multipart sessions, manifests, and visibility.
Partitioning / sharding
Hash keys across storage nodes and erasure-code chunks across failure domains; partition metadata by bucket/tenant + key hash.
Indexes
Unique bucket+key+version, uploads by owner/state/expiry, content digest, and lifecycle by class/time.
Replication + consistency
A version is visible only after every part verifies and the manifest commits. Bytes replicate/erasure-code independently.
Cache, queue + recovery
Signed URLs enable direct transfer; finalize events trigger scan/rendition/lifecycle; abandoned multipart uploads expire.
Capacity math
Estimate object count, p50/p99 size, ingress/egress Gbps, multipart concurrency, durability, retention, and repair bandwidth.
Alternative rejected
Database BLOBs suit small transactional files but make large-object backup/serving costly; separate bytes and metadata.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Objects, buckets, and keys

Objects are addressed by keys and replaced as whole values. Model folders as prefixes, not mutable directory trees.

Tie the mechanism back to Metadata database, Object/CDN index, and the stated object size distribution · concurrency · part size · retention · durability · egress envelope.
Direct upload

Issue a short-lived signed URL so the client transfers bytes without consuming application-server bandwidth.

Tie the mechanism back to Metadata database, Object/CDN index, and the stated object size distribution · concurrency · part size · retention · durability · egress envelope.
Multipart transfer

Split large files into checksummed parts for parallel upload, retry, resume, and final manifest assembly.

Tie the mechanism back to Metadata database, Object/CDN index, and the stated object size distribution · concurrency · part size · retention · durability · egress envelope.
05

Failure pressure test

Show detection, containment, recovery, and evidence.

The topic-specific correctness risk

Orphaned chunks and prematurely committed metadata require finalize and garbage-collection protocols.

Track failed promises at Metadata database and Object/CDN index.
Object event queue or Scan/lifecycle workers falls behind

Bound admission, scale on oldest-work age, retry with jitter, and isolate poison work before lag becomes unbounded.

Oldest event age · retry rate · dead-letter volume · projection freshness
Metadata database is slow or unavailable

Apply a deadline, preserve retry identity, fail over only within the stated consistency model, and reconcile any ambiguous result.

Commit p99 · timeout rate · replication lag · recovery time
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution

Read the solid request path first, stop at the source of truth, then follow the dashed event path into workers and rebuildable read models. Every arrow names a contract you should be ready to defend.

  1. 01

    Client — Requests an upload session Define the output contract before moving to the next owner.

  2. 02

    Metadata service — Allocates identity Define the output contract before moving to the next owner.

  3. 03

    Signed URL — Authorizes direct transfer Define the output contract before moving to the next owner.

  4. 04

    Chunk store — Persists parts Define the output contract before moving to the next owner.

  5. 05

    Verifier — Checks checksums Define the output contract before moving to the next owner.

  6. 06

    Lifecycle worker — Replicates, tiers, or expires Confirm the result and emit the evidence needed to reconcile it.

02

Lesson spine

What you need to understand.

Large immutable bytes belong in object storage; metadata and workflow state belong in a database that can coordinate visibility.

01

Objects, buckets, and keys

Objects are addressed by keys and replaced as whole values. Model folders as prefixes, not mutable directory trees.

02

Direct upload

Issue a short-lived signed URL so the client transfers bytes without consuming application-server bandwidth.

03

Multipart transfer

Split large files into checksummed parts for parallel upload, retry, resume, and final manifest assembly.

04

Durability and availability

Replicate across failure domains, verify checksums, and distinguish losing bytes from temporarily failing to serve them.

05

Processing pipeline

Publish an upload event, create derived renditions asynchronously, and expose only assets whose metadata reached a valid state.

06

Security and lifecycle

Use least-privilege signed access, encryption, malware scanning, retention, archival classes, and cleanup of abandoned parts.

03

Before the boxes

Frame the decision.

Outcome

What must work

Separate immutable large objects from metadata and design multipart upload, durability, and lifecycle policies.

Scale

What changes the design

Object size distribution · concurrency · part size · retention · durability · egress

Boundary

What owns the truth

Identify the component that commits authoritative state, then separate synchronous confirmation from derived work.

Non-goal

What stays simple

Do not add global coordination, multi-region writes, or a specialized store until a requirement earns the complexity.

04

Decision table

Make the trade-offs explicit.

DecisionDefensible positionCost to acknowledge
Primary mechanismDirect upload removes application bandwidth; proxying offers control at much greater cost.The stronger guarantee usually adds coordination, latency, state, or operational work.
Sync vs. asyncKeep only correctness-critical confirmation synchronous. Move derived views, notifications, analytics, and cleanup behind a durable boundary.Async work needs idempotency, lag monitoring, replay, and a product definition for partial completion.
Simple vs. scaledBegin with one logical owner and a clear API. Partition or replicate only the resource proven to be the first bottleneck.Migration requires stable identities, versioned contracts, backfill, and a rollback path.
05

Failure review

Design the recovery path.

DetectBoundRetry safelyReconcileLearn

Topic-specific risk

Orphaned chunks and prematurely committed metadata require finalize and garbage-collection protocols.

Response

Persist enough identity and state to distinguish retry, resume, compensation, and operator repair.

Dependency timeout

A timeout is ambiguous: the remote side may have failed, succeeded, or still be running.

Response

Use deadlines, bounded backoff with jitter, idempotency keys, and a status or reconciliation path.

Overload or skew

Average capacity can look healthy while a tenant, key, partition, region, or expensive request saturates one owner.

Response

Expose queue depth and hot-key share, apply backpressure, isolate tenants, and degrade optional work before correctness.

06

Evidence + level bar

Prove the design can be operated.

Core signals

Health of the promise

Measure user-visible latency or freshness, correctness drift, saturation, retry volume, and time to recover. Alert on the failed promise—not only CPU.

Mid-level

Complete and clear

Finish the happy path, identify the state owner, choose reasonable building blocks, and explain one scale mechanism.

Senior

Trade-offs and failure

Separate read and write paths, define consistency, explain partitioning, and make duplicate or partial failure safe.

Staff+

Evolution and operations

Discuss multi-region boundaries, migration, tenant isolation, capacity, observability, and how the architecture changes over time.

07

Interview language

Open the deep dive with a claim.

“For Blob storage, the decision I want to make explicit is this: Direct upload removes application bandwidth; proxying offers control at much greater cost. I’ll trace the state-changing path first, show where the result becomes durable, then test the design against the highest-risk failure and our target scale.”

08 · Retrieval check

Can you defend it without the page?

  1. For Blob storage, where is the correctness boundary and which failure would you test first?
  2. Which component owns committed truth, and what event or response proves the commit?
  3. Where is the first scaling or coordination bottleneck under the stated envelope?
  4. What happens after an ambiguous timeout or duplicate operation?
  5. Which complexity would you remove at one hundredth of the scale?