03Building blocksWrite-ahead logging

03 · Building block

Write-ahead logging

Use append-first durability for recovery, replication, and change streams while managing compaction.
8 minConcept guideReference-informed · independently authored
01

Architecture map

See where the component sits in a real system.

Write-ahead logging · system architecture
Write-ahead logging system architecture. The log reaches durable quorum before pages change; replay reconstructs state after failure. Request path: Database clients to DB protocol layer to Leader execution to Durable WAL. Asynchronous path: Replication stream to Replica/recovery. Read path: Read engine to Buffer pool + SST/pages. External dependency: Backup archive.

Write

Database clients enters through DB protocol layer. Leader execution owns validation and commits the durable record to Durable WAL.

Propagate

Replication stream separates the committed write from background work. Replica/recovery can retry safely while it builds Buffer pool + SST/pages.

Read

Read engine serves from Buffer pool + SST/pages, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Backup archive as an explicit dependency.

Say this first: The log reaches durable quorum before pages change; replay reconstructs state after failure.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

The log reaches durable quorum before pages change; replay reconstructs state after failure.

INTERVIEW CONTRACT

Use append-first durability for recovery, replication, and change streams while managing compaction.

CAPACITY QUESTIONS TO QUANTIFY

Append rate · record size · fsync p99 · replica lag · recovery objective. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Database clients → DB protocol layer

    Transactional mutations enters over HTTPS / RPC. DB protocol layer handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    DB protocol layer → Leader execution → Durable WAL

    Leader execution receives the command, checks invariants and retry identity, then uses fsync / quorum LSN to update Durable WAL. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Leader execution → Replication stream → Replica/recovery → Buffer pool + SST/pages

    Leader execution emits publish after commit; Replica/recovery uses consume and redo to build Buffer pool + SST/pages. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    DB protocol layer → Read engine → Buffer pool + SST/pages / Durable WAL

    Read engine uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    Replica/recovery → Backup archive

    ship segments crosses into Backup archive. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
DB protocol layerSession + transactionIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Leader executionValidate + assign LSNWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Durable WALOrdered redo recordsAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Replication streamLSN-ordered bytesDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Replica/recoveryApply redo + checkpointReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Buffer pool + SST/pagesMaterialized database stateRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Read engineSnapshot + page lookupRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Backup archiveSegments + snapshotsExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
Sequential WAL segments live on durable local/NVMe storage replicated with Raft/Paxos; object storage archives segments/snapshots.
Partitioning / sharding
The database shards first by key range/hash; each shard has one ordered WAL leader and independent quorum.
Indexes
Sequential LSN, sparse segment offset, transaction commit table, and archive keys by shard/LSN range.
Replication + consistency
Acknowledge only after configured fsync/quorum LSN. Followers expose applied LSN so replica staleness is measurable.
Cache, queue + recovery
Group commit improves throughput; checkpoints bound replay; replicas/CDC checkpoint LSNs and replay idempotently.
Capacity math
Estimate log bytes/write, writes/sec, fsync p99, batch size, retention, lag, checkpoint duration, and recovery time.
Alternative rejected
Mutating pages before logging risks torn recovery; append-first durability creates an ordered reconstruction path.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Append before apply

Assign a log sequence number, append the record, force the required bytes to stable storage, then mutate memory and acknowledge.

Tie the mechanism back to Durable WAL, Buffer pool + SST/pages, and the stated append rate · record size · fsync p99 · replica lag · recovery objective envelope.
Crash recovery

Load the last checkpoint and replay later redo records idempotently until the durable end of the log.

Tie the mechanism back to Durable WAL, Buffer pool + SST/pages, and the stated append rate · record size · fsync p99 · replica lag · recovery objective envelope.
Redo and undo

Redo logs reconstruct committed changes; undo information helps roll back changes that were durable but not committed.

Tie the mechanism back to Durable WAL, Buffer pool + SST/pages, and the stated append rate · record size · fsync p99 · replica lag · recovery objective envelope.
05

Failure pressure test

Show detection, containment, recovery, and evidence.

The topic-specific correctness risk

A corrupt tail or replayed non-idempotent effect can turn recovery into more damage.

Track failed promises at Durable WAL and Buffer pool + SST/pages.
Replication stream or Replica/recovery falls behind

Bound admission, scale on oldest-work age, retry with jitter, and isolate poison work before lag becomes unbounded.

Oldest event age · retry rate · dead-letter volume · projection freshness
Durable WAL is slow or unavailable

Apply a deadline, preserve retry identity, fail over only within the stated consistency model, and reconcile any ambiguous result.

Commit p99 · timeout rate · replication lag · recovery time
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution

Read the solid request path first, stop at the source of truth, then follow the dashed event path into workers and rebuildable read models. Every arrow names a contract you should be ready to defend.

  1. 01

    Command — Carries operation identity Define the output contract before moving to the next owner.

  2. 02

    Log writer — Appends and fsyncs Define the output contract before moving to the next owner.

  3. 03

    Commit index — Marks durable order Define the output contract before moving to the next owner.

  4. 04

    State machine — Applies records Define the output contract before moving to the next owner.

  5. 05

    Replica — Replays entries Define the output contract before moving to the next owner.

  6. 06

    Compactor — Snapshots and reclaims Confirm the result and emit the evidence needed to reconcile it.

02

Lesson spine

What you need to understand.

Write-ahead logging makes a change durable before slower data structures are updated, enabling crash recovery and ordered replication.

01

Append before apply

Assign a log sequence number, append the record, force the required bytes to stable storage, then mutate memory and acknowledge.

02

Crash recovery

Load the last checkpoint and replay later redo records idempotently until the durable end of the log.

03

Redo and undo

Redo logs reconstruct committed changes; undo information helps roll back changes that were durable but not committed.

04

Checkpoints

Flush dirty state, record the safe log position, and truncate only data no longer needed for recovery or replicas.

05

Replication and CDC

The ordered log can feed followers and downstream projections, but consumers still need versioning and replay-safe effects.

06

Operational limits

Batch fsync for throughput, monitor flush latency and log growth, and test torn writes and incomplete records.

03

Before the boxes

Frame the decision.

Outcome

What must work

Use append-first durability for recovery, replication, and change streams while managing compaction.

Scale

What changes the design

Append rate · record size · fsync p99 · replica lag · recovery objective

Boundary

What owns the truth

Identify the component that commits authoritative state, then separate synchronous confirmation from derived work.

Non-goal

What stays simple

Do not add global coordination, multi-region writes, or a specialized store until a requirement earns the complexity.

04

Decision table

Make the trade-offs explicit.

DecisionDefensible positionCost to acknowledge
Primary mechanismGroup commit shares durable-write cost; per-record fsync minimizes acknowledged-loss risk.The stronger guarantee usually adds coordination, latency, state, or operational work.
Sync vs. asyncKeep only correctness-critical confirmation synchronous. Move derived views, notifications, analytics, and cleanup behind a durable boundary.Async work needs idempotency, lag monitoring, replay, and a product definition for partial completion.
Simple vs. scaledBegin with one logical owner and a clear API. Partition or replicate only the resource proven to be the first bottleneck.Migration requires stable identities, versioned contracts, backfill, and a rollback path.
05

Failure review

Design the recovery path.

DetectBoundRetry safelyReconcileLearn

Topic-specific risk

A corrupt tail or replayed non-idempotent effect can turn recovery into more damage.

Response

Persist enough identity and state to distinguish retry, resume, compensation, and operator repair.

Dependency timeout

A timeout is ambiguous: the remote side may have failed, succeeded, or still be running.

Response

Use deadlines, bounded backoff with jitter, idempotency keys, and a status or reconciliation path.

Overload or skew

Average capacity can look healthy while a tenant, key, partition, region, or expensive request saturates one owner.

Response

Expose queue depth and hot-key share, apply backpressure, isolate tenants, and degrade optional work before correctness.

06

Evidence + level bar

Prove the design can be operated.

Core signals

Health of the promise

Measure user-visible latency or freshness, correctness drift, saturation, retry volume, and time to recover. Alert on the failed promise—not only CPU.

Mid-level

Complete and clear

Finish the happy path, identify the state owner, choose reasonable building blocks, and explain one scale mechanism.

Senior

Trade-offs and failure

Separate read and write paths, define consistency, explain partitioning, and make duplicate or partial failure safe.

Staff+

Evolution and operations

Discuss multi-region boundaries, migration, tenant isolation, capacity, observability, and how the architecture changes over time.

07

Interview language

Open the deep dive with a claim.

“For Write-ahead logging, the decision I want to make explicit is this: Group commit shares durable-write cost; per-record fsync minimizes acknowledged-loss risk. I’ll trace the state-changing path first, show where the result becomes durable, then test the design against the highest-risk failure and our target scale.”

08 · Retrieval check

Can you defend it without the page?

  1. For Write-ahead logging, where is the correctness boundary and which failure would you test first?
  2. Which component owns committed truth, and what event or response proves the commit?
  3. Where is the first scaling or coordination bottleneck under the stated envelope?
  4. What happens after an ambiguous timeout or duplicate operation?
  5. Which complexity would you remove at one hundredth of the scale?