03Building blocksMessage queues

03 · Building block

Message queues

Decouple work with durable delivery, consumer groups, retries, dead letters, ordering, and backpressure.
8 minConcept guideReference-informed · independently authored
01

Architecture map

See where the component sits in a real system.

Message queues · system architecture
Message queues system architecture. The broker durably owns unacknowledged work; consumers own effects only after their idempotent boundary. Request path: Producers to Broker ingress to Partition leader to Replicated broker log. Asynchronous path: Topic partitions to Consumer group. Read path: Consumer fetch to Offsets + retry state. External dependency: Effect database/API.

Write

Producers enters through Broker ingress. Partition leader owns validation and commits the durable record to Replicated broker log.

Propagate

Topic partitions separates the committed write from background work. Consumer group can retry safely while it builds Offsets + retry state.

Read

Consumer fetch serves from Offsets + retry state, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Effect database/API as an explicit dependency.

Say this first: The broker durably owns unacknowledged work; consumers own effects only after their idempotent boundary.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

The broker durably owns unacknowledged work; consumers own effects only after their idempotent boundary.

INTERVIEW CONTRACT

Decouple work with durable delivery, consumer groups, retries, dead letters, ordering, and backpressure.

CAPACITY QUESTIONS TO QUANTIFY

Publish rate · consumer rate · end-to-end lag · retry age · partition skew. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Producers → Broker ingress

    Jobs or domain events enters over HTTPS / RPC. Broker ingress handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Broker ingress → Partition leader → Replicated broker log

    Partition leader receives the command, checks invariants and retry identity, then uses replicated append to update Replicated broker log. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Partition leader → Topic partitions → Consumer group → Offsets + retry state

    Partition leader emits publish after commit; Consumer group uses consume and ack after commit to build Offsets + retry state. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Broker ingress → Consumer fetch → Offsets + retry state / Replicated broker log

    Consumer fetch uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    Consumer group → Effect database/API

    apply effect crosses into Effect database/API. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Broker ingressAuth + partition keyIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Partition leaderAppend + replicateWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Replicated broker logDurable ordered recordsAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Topic partitionsParallelism + orderingDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Consumer groupProcess + ack/commitReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Offsets + retry stateProgress, backoff, DLQRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Consumer fetchOffset / visibility leaseRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Effect database/APIIdempotent side effectExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
Kafka/Pulsar fits replayable ordered logs, RabbitMQ routing-heavy queues, and SQS/Pub/Sub managed at-least-once work.
Partitioning / sharding
Partition by the smallest key needing order—aggregate, customer, job, or device—so unrelated keys process in parallel.
Indexes
Append by partition/offset, group progress by partition, retry by next_visible_at, and compacted keys when latest state matters.
Replication + consistency
RF=3 with quorum acknowledgement for durable topics. Default to at least once; commit progress only after idempotent effects.
Cache, queue + recovery
Use jittered retry, bounded attempts, DLQ/repair, schema versions, lag alarms, and fenced consumer-group rebalance.
Capacity math
Estimate messages/sec, bytes/sec, retention, partition throughput, processing time, lag SLA, replay window, and poison rate.
Alternative rejected
A queue per consumer preserves order but destroys scale; partitions make the order/parallelism boundary explicit.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Queue vs topic

A queue distributes work among consumers. A topic lets independent consumer groups react to the same event.

Tie the mechanism back to Replicated broker log, Offsets + retry state, and the stated publish rate · consumer rate · end-to-end lag · retry age · partition skew envelope.
Delivery semantics

At-most-once may lose work, at-least-once may repeat work, and practical exactly-once effects require idempotent processing and an atomic boundary.

Tie the mechanism back to Replicated broker log, Offsets + retry state, and the stated publish rate · consumer rate · end-to-end lag · retry age · partition skew envelope.
Ordering

Preserve order only within a key or partition that needs it; global order destroys parallelism.

Tie the mechanism back to Replicated broker log, Offsets + retry state, and the stated publish rate · consumer rate · end-to-end lag · retry age · partition skew envelope.
05

Failure pressure test

Show detection, containment, recovery, and evidence.

The topic-specific correctness risk

Poison messages and slow consumers can exhaust retention and turn lag into data loss.

Track failed promises at Replicated broker log and Offsets + retry state.
Topic partitions or Consumer group falls behind

Bound admission, scale on oldest-work age, retry with jitter, and isolate poison work before lag becomes unbounded.

Oldest event age · retry rate · dead-letter volume · projection freshness
Replicated broker log is slow or unavailable

Apply a deadline, preserve retry identity, fail over only within the stated consistency model, and reconcile any ambiguous result.

Commit p99 · timeout rate · replication lag · recovery time
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution

Read the solid request path first, stop at the source of truth, then follow the dashed event path into workers and rebuildable read models. Every arrow names a contract you should be ready to defend.

  1. 01

    Producer — Publishes an event Define the output contract before moving to the next owner.

  2. 02

    Broker — Persists and partitions Define the output contract before moving to the next owner.

  3. 03

    Consumer group — Shares ordered work Define the output contract before moving to the next owner.

  4. 04

    Handler — Applies an idempotent effect Define the output contract before moving to the next owner.

  5. 05

    Retry lane — Delays transient errors Define the output contract before moving to the next owner.

  6. 06

    Dead letter queue — Quarantines poison messages Confirm the result and emit the evidence needed to reconcile it.

02

Lesson spine

What you need to understand.

Queues create a durable time boundary between producers and consumers; that boundary changes failure, load, and ownership.

01

Queue vs topic

A queue distributes work among consumers. A topic lets independent consumer groups react to the same event.

02

Delivery semantics

At-most-once may lose work, at-least-once may repeat work, and practical exactly-once effects require idempotent processing and an atomic boundary.

03

Ordering

Preserve order only within a key or partition that needs it; global order destroys parallelism.

04

Backpressure

Track lag, queue age, and worker saturation; shed optional production or scale consumers before unbounded delay.

05

Retries and dead letters

Use bounded exponential backoff, classify permanent failures, and make poison messages visible for repair.

06

Technology fit

Kafka favors replayable ordered logs, RabbitMQ rich routing and work queues, and managed queues operational simplicity.

03

Before the boxes

Frame the decision.

Outcome

What must work

Decouple work with durable delivery, consumer groups, retries, dead letters, ordering, and backpressure.

Scale

What changes the design

Publish rate · consumer rate · end-to-end lag · retry age · partition skew

Boundary

What owns the truth

Identify the component that commits authoritative state, then separate synchronous confirmation from derived work.

Non-goal

What stays simple

Do not add global coordination, multi-region writes, or a specialized store until a requirement earns the complexity.

04

Decision table

Make the trade-offs explicit.

DecisionDefensible positionCost to acknowledge
Primary mechanismAt-least-once delivery plus idempotent consumers is usually the honest contract.The stronger guarantee usually adds coordination, latency, state, or operational work.
Sync vs. asyncKeep only correctness-critical confirmation synchronous. Move derived views, notifications, analytics, and cleanup behind a durable boundary.Async work needs idempotency, lag monitoring, replay, and a product definition for partial completion.
Simple vs. scaledBegin with one logical owner and a clear API. Partition or replicate only the resource proven to be the first bottleneck.Migration requires stable identities, versioned contracts, backfill, and a rollback path.
05

Failure review

Design the recovery path.

DetectBoundRetry safelyReconcileLearn

Topic-specific risk

Poison messages and slow consumers can exhaust retention and turn lag into data loss.

Response

Persist enough identity and state to distinguish retry, resume, compensation, and operator repair.

Dependency timeout

A timeout is ambiguous: the remote side may have failed, succeeded, or still be running.

Response

Use deadlines, bounded backoff with jitter, idempotency keys, and a status or reconciliation path.

Overload or skew

Average capacity can look healthy while a tenant, key, partition, region, or expensive request saturates one owner.

Response

Expose queue depth and hot-key share, apply backpressure, isolate tenants, and degrade optional work before correctness.

06

Evidence + level bar

Prove the design can be operated.

Core signals

Health of the promise

Measure user-visible latency or freshness, correctness drift, saturation, retry volume, and time to recover. Alert on the failed promise—not only CPU.

Mid-level

Complete and clear

Finish the happy path, identify the state owner, choose reasonable building blocks, and explain one scale mechanism.

Senior

Trade-offs and failure

Separate read and write paths, define consistency, explain partitioning, and make duplicate or partial failure safe.

Staff+

Evolution and operations

Discuss multi-region boundaries, migration, tenant isolation, capacity, observability, and how the architecture changes over time.

07

Interview language

Open the deep dive with a claim.

“For Message queues, the decision I want to make explicit is this: At-least-once delivery plus idempotent consumers is usually the honest contract. I’ll trace the state-changing path first, show where the result becomes durable, then test the design against the highest-risk failure and our target scale.”

08 · Retrieval check

Can you defend it without the page?

  1. For Message queues, where is the correctness boundary and which failure would you test first?
  2. Which component owns committed truth, and what event or response proves the commit?
  3. Where is the first scaling or coordination bottleneck under the stated envelope?
  4. What happens after an ambiguous timeout or duplicate operation?
  5. Which complexity would you remove at one hundredth of the scale?