03Building blocksMessage queues
03 · Building block
Message queues
Decouple work with durable delivery, consumer groups, retries, dead letters, ordering, and backpressure.Architecture map
See where the component sits in a real system.

Write
Producers enters through Broker ingress. Partition leader owns validation and commits the durable record to Replicated broker log.
Propagate
Topic partitions separates the committed write from background work. Consumer group can retry safely while it builds Offsets + retry state.
Read
Consumer fetch serves from Offsets + retry state, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Effect database/API as an explicit dependency.
Say this first: The broker durably owns unacknowledged work; consumers own effects only after their idempotent boundary.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
The broker durably owns unacknowledged work; consumers own effects only after their idempotent boundary.
Decouple work with durable delivery, consumer groups, retries, dead letters, ordering, and backpressure.
Publish rate · consumer rate · end-to-end lag · retry age · partition skew. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
Producers → Broker ingressJobs or domain events enters over HTTPS / RPC. Broker ingress handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Broker ingress → Partition leader → Replicated broker logPartition leader receives the command, checks invariants and retry identity, then uses replicated append to update Replicated broker log. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Partition leader → Topic partitions → Consumer group → Offsets + retry statePartition leader emits publish after commit; Consumer group uses consume and ack after commit to build Offsets + retry state. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Broker ingress → Consumer fetch → Offsets + retry state / Replicated broker logConsumer fetch uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
Consumer group → Effect database/APIapply effect crosses into Effect database/API. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Broker ingressAuth + partition key | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Partition leaderAppend + replicate | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Replicated broker logDurable ordered records | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| Topic partitionsParallelism + ordering | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| Consumer groupProcess + ack/commit | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Offsets + retry stateProgress, backoff, DLQ | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Consumer fetchOffset / visibility lease | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Effect database/APIIdempotent side effect | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- Kafka/Pulsar fits replayable ordered logs, RabbitMQ routing-heavy queues, and SQS/Pub/Sub managed at-least-once work.
- Partitioning / sharding
- Partition by the smallest key needing order—aggregate, customer, job, or device—so unrelated keys process in parallel.
- Indexes
- Append by partition/offset, group progress by partition, retry by next_visible_at, and compacted keys when latest state matters.
- Replication + consistency
- RF=3 with quorum acknowledgement for durable topics. Default to at least once; commit progress only after idempotent effects.
- Cache, queue + recovery
- Use jittered retry, bounded attempts, DLQ/repair, schema versions, lag alarms, and fenced consumer-group rebalance.
- Capacity math
- Estimate messages/sec, bytes/sec, retention, partition throughput, processing time, lag SLA, replay window, and poison rate.
- Alternative rejected
- A queue per consumer preserves order but destroys scale; partitions make the order/parallelism boundary explicit.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
Queue vs topic
A queue distributes work among consumers. A topic lets independent consumer groups react to the same event.
Tie the mechanism back to Replicated broker log, Offsets + retry state, and the stated publish rate · consumer rate · end-to-end lag · retry age · partition skew envelope.Delivery semantics
At-most-once may lose work, at-least-once may repeat work, and practical exactly-once effects require idempotent processing and an atomic boundary.
Tie the mechanism back to Replicated broker log, Offsets + retry state, and the stated publish rate · consumer rate · end-to-end lag · retry age · partition skew envelope.Ordering
Preserve order only within a key or partition that needs it; global order destroys parallelism.
Tie the mechanism back to Replicated broker log, Offsets + retry state, and the stated publish rate · consumer rate · end-to-end lag · retry age · partition skew envelope.Failure pressure test
Show detection, containment, recovery, and evidence.
The topic-specific correctness risk
Poison messages and slow consumers can exhaust retention and turn lag into data loss.
Track failed promises at Replicated broker log and Offsets + retry state.Topic partitions or Consumer group falls behind
Bound admission, scale on oldest-work age, retry with jitter, and isolate poison work before lag becomes unbounded.
Oldest event age · retry rate · dead-letter volume · projection freshnessReplicated broker log is slow or unavailable
Apply a deadline, preserve retry identity, fail over only within the stated consistency model, and reconcile any ambiguous result.
Commit p99 · timeout rate · replication lag · recovery time- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
Read the solid request path first, stop at the source of truth, then follow the dashed event path into workers and rebuildable read models. Every arrow names a contract you should be ready to defend.
- 01
Producer — Publishes an event Define the output contract before moving to the next owner.
- 02
Broker — Persists and partitions Define the output contract before moving to the next owner.
- 03
Consumer group — Shares ordered work Define the output contract before moving to the next owner.
- 04
Handler — Applies an idempotent effect Define the output contract before moving to the next owner.
- 05
Retry lane — Delays transient errors Define the output contract before moving to the next owner.
- 06
Dead letter queue — Quarantines poison messages Confirm the result and emit the evidence needed to reconcile it.
Lesson spine
What you need to understand.
Queues create a durable time boundary between producers and consumers; that boundary changes failure, load, and ownership.
Queue vs topic
A queue distributes work among consumers. A topic lets independent consumer groups react to the same event.
Delivery semantics
At-most-once may lose work, at-least-once may repeat work, and practical exactly-once effects require idempotent processing and an atomic boundary.
Ordering
Preserve order only within a key or partition that needs it; global order destroys parallelism.
Backpressure
Track lag, queue age, and worker saturation; shed optional production or scale consumers before unbounded delay.
Retries and dead letters
Use bounded exponential backoff, classify permanent failures, and make poison messages visible for repair.
Technology fit
Kafka favors replayable ordered logs, RabbitMQ rich routing and work queues, and managed queues operational simplicity.
Before the boxes
Frame the decision.
What must work
Decouple work with durable delivery, consumer groups, retries, dead letters, ordering, and backpressure.
What changes the design
Publish rate · consumer rate · end-to-end lag · retry age · partition skew
What owns the truth
Identify the component that commits authoritative state, then separate synchronous confirmation from derived work.
What stays simple
Do not add global coordination, multi-region writes, or a specialized store until a requirement earns the complexity.
Decision table
Make the trade-offs explicit.
| Decision | Defensible position | Cost to acknowledge |
|---|---|---|
| Primary mechanism | At-least-once delivery plus idempotent consumers is usually the honest contract. | The stronger guarantee usually adds coordination, latency, state, or operational work. |
| Sync vs. async | Keep only correctness-critical confirmation synchronous. Move derived views, notifications, analytics, and cleanup behind a durable boundary. | Async work needs idempotency, lag monitoring, replay, and a product definition for partial completion. |
| Simple vs. scaled | Begin with one logical owner and a clear API. Partition or replicate only the resource proven to be the first bottleneck. | Migration requires stable identities, versioned contracts, backfill, and a rollback path. |
Failure review
Design the recovery path.
Topic-specific risk
Poison messages and slow consumers can exhaust retention and turn lag into data loss.
ResponsePersist enough identity and state to distinguish retry, resume, compensation, and operator repair.
Dependency timeout
A timeout is ambiguous: the remote side may have failed, succeeded, or still be running.
ResponseUse deadlines, bounded backoff with jitter, idempotency keys, and a status or reconciliation path.
Overload or skew
Average capacity can look healthy while a tenant, key, partition, region, or expensive request saturates one owner.
ResponseExpose queue depth and hot-key share, apply backpressure, isolate tenants, and degrade optional work before correctness.
Evidence + level bar
Prove the design can be operated.
Health of the promise
Measure user-visible latency or freshness, correctness drift, saturation, retry volume, and time to recover. Alert on the failed promise—not only CPU.
Complete and clear
Finish the happy path, identify the state owner, choose reasonable building blocks, and explain one scale mechanism.
Trade-offs and failure
Separate read and write paths, define consistency, explain partitioning, and make duplicate or partial failure safe.
Evolution and operations
Discuss multi-region boundaries, migration, tenant isolation, capacity, observability, and how the architecture changes over time.
Interview language
Open the deep dive with a claim.
“For Message queues, the decision I want to make explicit is this: At-least-once delivery plus idempotent consumers is usually the honest contract. I’ll trace the state-changing path first, show where the result becomes durable, then test the design against the highest-risk failure and our target scale.”
08 · Retrieval check
Can you defend it without the page?
- For Message queues, where is the correctness boundary and which failure would you test first?
- Which component owns committed truth, and what event or response proves the commit?
- Where is the first scaling or coordination bottleneck under the stated envelope?
- What happens after an ambiguous timeout or duplicate operation?
- Which complexity would you remove at one hundredth of the scale?