03Building blocksLoad balancers

03 · Building block

Load balancers

Distribute traffic, detect unhealthy instances, preserve affinity only when necessary, and remove single points of failure.
8 minConcept guideReference-informed · independently authored
01

Architecture map

See where the component sits in a real system.

Load balancers · system architecture
Load balancers system architecture. A stateless traffic tier routes only to a health-checked backend set and drains ownership safely. Request path: Internet clients to DNS / edge LB to L7 load balancer to Service registry. Asynchronous path: Health observations to Health controller. Read path: Backend selection to Healthy backend set. External dependency: Application fleet.

Write

Internet clients enters through DNS / edge LB. L7 load balancer owns validation and commits the durable record to Service registry.

Propagate

Health observations separates the committed write from background work. Health controller can retry safely while it builds Healthy backend set.

Read

Backend selection serves from Healthy backend set, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Application fleet as an explicit dependency.

Say this first: A stateless traffic tier routes only to a health-checked backend set and drains ownership safely.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

A stateless traffic tier routes only to a health-checked backend set and drains ownership safely.

INTERVIEW CONTRACT

Distribute traffic, detect unhealthy instances, preserve affinity only when necessary, and remove single points of failure.

CAPACITY QUESTIONS TO QUANTIFY

Peak connections · QPS · TLS cost · cross-zone traffic · failover headroom. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Internet clients → DNS / edge LB

    HTTPS requests enters over HTTPS / RPC. DNS / edge LB handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    DNS / edge LB → L7 load balancer → Service registry

    L7 load balancer receives the command, checks invariants and retry identity, then uses commit to update Service registry. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    L7 load balancer → Health observations → Health controller → Healthy backend set

    L7 load balancer emits publish after commit; Health controller uses consume and publish endpoint set to build Healthy backend set. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    DNS / edge LB → Backend selection → Healthy backend set / Service registry

    Backend selection uses optimized read for the common, read-optimized path and watch registry when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    L7 load balancer → Application fleet

    proxied request crosses into Application fleet. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
DNS / edge LBRegion + anycast routingIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
L7 load balancerTLS, routing, affinityWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Service registryDesired endpoints + weightsAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Health observationsActive + passive checksDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Health controllerEject, recover, drainReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Healthy backend setLocal routing snapshotRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Backend selectionLeast-load / hash / RRRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Application fleetStateless service instancesExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
Stateless L4/L7 balancers use etcd/Consul or managed registry truth and local in-memory healthy endpoint snapshots.
Partitioning / sharding
Partition traffic by region/service; use rendezvous hashing only when affinity/cache locality is a stated requirement.
Indexes
Registry by service/zone/state; local endpoint array/ring; health windows by endpoint; route trie for L7.
Replication + consistency
Config/registry versions are strong; health is freshness-bounded. Last-known-good snapshots keep data plane alive during control loss.
Cache, queue + recovery
Active/passive health, slow start, bounded ejection, draining, retry budgets, and zone-aware routing form the failure policy.
Capacity math
Estimate QPS, connections, bytes/sec, TLS handshakes, backend count, zone loss, and longest request for draining.
Alternative rejected
Round robin is correct for uniform work; least-load/EWMA/hashing needs variable cost or affinity evidence.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Layer 4 vs layer 7

L4 routes connections with low overhead; L7 understands HTTP, hosts, paths, headers, and application policy.

Tie the mechanism back to Service registry, Healthy backend set, and the stated peak connections · qps · tls cost · cross-zone traffic · failover headroom envelope.
Algorithms

Round robin suits similar requests; least connections or response time helps variable work; weighted policies reflect unequal capacity.

Tie the mechanism back to Service registry, Healthy backend set, and the stated peak connections · qps · tls cost · cross-zone traffic · failover headroom envelope.
Health checks

Combine active probes with passive error signals, slow start, draining, and bounded ejection so deployment churn does not become an outage.

Tie the mechanism back to Service registry, Healthy backend set, and the stated peak connections · qps · tls cost · cross-zone traffic · failover headroom envelope.
05

Failure pressure test

Show detection, containment, recovery, and evidence.

The topic-specific correctness risk

Slow or flapping instances can pass shallow health checks and absorb failing traffic.

Track failed promises at Service registry and Healthy backend set.
Health observations or Health controller falls behind

Bound admission, scale on oldest-work age, retry with jitter, and isolate poison work before lag becomes unbounded.

Oldest event age · retry rate · dead-letter volume · projection freshness
Service registry is slow or unavailable

Apply a deadline, preserve retry identity, fail over only within the stated consistency model, and reconcile any ambiguous result.

Commit p99 · timeout rate · replication lag · recovery time
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution

Read the solid request path first, stop at the source of truth, then follow the dashed event path into workers and rebuildable read models. Every arrow names a contract you should be ready to defend.

  1. 01

    Client — Resolves the service Define the output contract before moving to the next owner.

  2. 02

    Global router — Chooses a region Define the output contract before moving to the next owner.

  3. 03

    L7 balancer — Routes requests Define the output contract before moving to the next owner.

  4. 04

    Service pool — Handles stateless work Define the output contract before moving to the next owner.

  5. 05

    Health system — Removes bad instances Define the output contract before moving to the next owner.

  6. 06

    Autoscaler — Adjusts capacity Confirm the result and emit the evidence needed to reconcile it.

02

Lesson spine

What you need to understand.

Load balancers distribute work and remove unhealthy capacity while preserving the routing information the application actually needs.

01

Layer 4 vs layer 7

L4 routes connections with low overhead; L7 understands HTTP, hosts, paths, headers, and application policy.

02

Algorithms

Round robin suits similar requests; least connections or response time helps variable work; weighted policies reflect unequal capacity.

03

Health checks

Combine active probes with passive error signals, slow start, draining, and bounded ejection so deployment churn does not become an outage.

04

Affinity

Use sticky routing only for state that cannot yet move; prefer external session state so a failed instance is replaceable.

05

Global routing

Use latency, geography, regulation, and regional health to select an entry region, then balance locally.

06

Zero-downtime changes

Register new capacity, verify readiness, shift traffic gradually, drain old connections, and retain rollback.

03

Before the boxes

Frame the decision.

Outcome

What must work

Distribute traffic, detect unhealthy instances, preserve affinity only when necessary, and remove single points of failure.

Scale

What changes the design

Peak connections · QPS · TLS cost · cross-zone traffic · failover headroom

Boundary

What owns the truth

Identify the component that commits authoritative state, then separate synchronous confirmation from derived work.

Non-goal

What stays simple

Do not add global coordination, multi-region writes, or a specialized store until a requirement earns the complexity.

04

Decision table

Make the trade-offs explicit.

DecisionDefensible positionCost to acknowledge
Primary mechanismLayer 4 minimizes overhead; Layer 7 earns application-aware routing and policy.The stronger guarantee usually adds coordination, latency, state, or operational work.
Sync vs. asyncKeep only correctness-critical confirmation synchronous. Move derived views, notifications, analytics, and cleanup behind a durable boundary.Async work needs idempotency, lag monitoring, replay, and a product definition for partial completion.
Simple vs. scaledBegin with one logical owner and a clear API. Partition or replicate only the resource proven to be the first bottleneck.Migration requires stable identities, versioned contracts, backfill, and a rollback path.
05

Failure review

Design the recovery path.

DetectBoundRetry safelyReconcileLearn

Topic-specific risk

Slow or flapping instances can pass shallow health checks and absorb failing traffic.

Response

Persist enough identity and state to distinguish retry, resume, compensation, and operator repair.

Dependency timeout

A timeout is ambiguous: the remote side may have failed, succeeded, or still be running.

Response

Use deadlines, bounded backoff with jitter, idempotency keys, and a status or reconciliation path.

Overload or skew

Average capacity can look healthy while a tenant, key, partition, region, or expensive request saturates one owner.

Response

Expose queue depth and hot-key share, apply backpressure, isolate tenants, and degrade optional work before correctness.

06

Evidence + level bar

Prove the design can be operated.

Core signals

Health of the promise

Measure user-visible latency or freshness, correctness drift, saturation, retry volume, and time to recover. Alert on the failed promise—not only CPU.

Mid-level

Complete and clear

Finish the happy path, identify the state owner, choose reasonable building blocks, and explain one scale mechanism.

Senior

Trade-offs and failure

Separate read and write paths, define consistency, explain partitioning, and make duplicate or partial failure safe.

Staff+

Evolution and operations

Discuss multi-region boundaries, migration, tenant isolation, capacity, observability, and how the architecture changes over time.

07

Interview language

Open the deep dive with a claim.

“For Load balancers, the decision I want to make explicit is this: Layer 4 minimizes overhead; Layer 7 earns application-aware routing and policy. I’ll trace the state-changing path first, show where the result becomes durable, then test the design against the highest-risk failure and our target scale.”

08 · Retrieval check

Can you defend it without the page?

  1. For Load balancers, where is the correctness boundary and which failure would you test first?
  2. Which component owns committed truth, and what event or response proves the commit?
  3. Where is the first scaling or coordination bottleneck under the stated envelope?
  4. What happens after an ambiguous timeout or duplicate operation?
  5. Which complexity would you remove at one hundredth of the scale?