04QuestionsWebhook delivery system
04 · Worked prompt
Webhook delivery system
Design subscriptions, signing, durable delivery, exponential retry, idempotency, and tenant isolation.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design subscriptions, signing, durable delivery, exponential retry, idempotency, and tenant isolation.

Write
Event producers enters through Webhook API. Delivery planner owns validation and commits the durable record to Delivery DB.
Propagate
Due-time queues separates the committed write from background work. Delivery workers can retry safely while it builds Endpoint health view.
Read
Delivery status API serves from Endpoint health view, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Subscriber endpoints as an explicit dependency.
Say this first: Each delivery attempt is durable; signing, backoff, and endpoint isolation make at-least-once delivery operable.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
Each delivery attempt is durable; signing, backoff, and endpoint isolation make at-least-once delivery operable.
Design subscriptions, signing, durable delivery, exponential retry, idempotency, and tenant isolation.
At-least-once · seconds first attempt · days retry · strict tenant quotas. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
Event producers → Webhook APIPublish tenant events enters over HTTPS / RPC. Webhook API handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Webhook API → Delivery planner → Delivery DBDelivery planner receives the command, checks invariants and retry identity, then uses create deliveries to update Delivery DB. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Delivery planner → Due-time queues → Delivery workers → Endpoint health viewDelivery planner emits publish after commit; Delivery workers uses consume and record outcome to build Endpoint health view. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Webhook API → Delivery status API → Endpoint health view / Delivery DBDelivery status API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
Delivery workers → Subscriber endpointssigned POST crosses into Subscriber endpoints. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Webhook APIAuth + subscription version | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Delivery plannerMatch subscriptions + due time | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Delivery DBEvent, endpoint, attempt state | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| Due-time queuesTenant + next-attempt bucket | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| Delivery workersSign + bounded HTTP send | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Endpoint health viewLatency, failures, pause | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Delivery status APIAttempts + terminal reason | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Subscriber endpointsIdempotent HTTPS receivers | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- PostgreSQL stores subscriptions/deliveries/attempts; Kafka/SQS stores due work; object storage retains large payloads.
- Partitioning / sharding
- Partition by tenant/endpoint + time bucket; preserve per-endpoint order only if required and isolate huge tenants.
- Indexes
- Unique (event_id, endpoint_id), due (state, next_attempt, shard), attempts by delivery/time, and endpoint failure streak.
- Replication + consistency
- Producer acceptance requires durable delivery creation. Delivery is at least once with stable event IDs and conditional local transitions.
- Cache, queue + recovery
- Queues carry delivery IDs; use visibility leases, exponential backoff+jitter, per-endpoint concurrency, circuit breaking, and DLQs.
- Capacity math
- Estimate events/sec × subscribers, payload size, endpoint latency, retry amplification, retention, and worst endpoint.
- Alternative rejected
- Synchronous fan-out couples producer availability to subscribers; durable delivery records isolate slow and failing endpoints.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
Retry policy
Use exponential backoff with jitter, maximum age, and Retry-After support
Retry forever is a denial-of-service featureOrdering
Offer per-subscription ordered mode only when required and expose head-of-line cost
Global ordering is unnecessaryTenant isolation
Partition quotas by tenant and endpoint, cap concurrent connections, and pause unhealthy targets
One customer must not consume the delivery fleetFailure pressure test
Show detection, containment, recovery, and evidence.
Endpoint times out
Record ambiguous attempt and retry with same event identity
timeout rate and oldest pendingSecret rotation
Sign with versioned active secret and allow bounded overlap
verification failures by secret versionDNS or certificate failure
Back off endpoint-wide and surface clear health state
paused endpoints- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize register event subscriptions and deliver signed events at least once.”
- Scale
“The design changes around at-least-once · seconds first attempt · days retry · strict tenant quotas.”
- Decision
“Per-endpoint order protects consumers but allows one poison event to block the stream.”
- Risk
“The first failure I want to pressure-test is: Slow receivers and unbounded retries can consume the entire fleet.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Register event subscriptions
- Deliver signed events at least once
- Retry transient failures for days
- Protect the fleet from slow or broken tenant endpoints
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 4 interfaces
Core entities
subscription_id, tenant_id, endpoint, secret_version, filtersOwner: Subscription serviceevent_id, type, payload_ref, occurred_atOwner: Event storedelivery_id, event_id, subscription_id, attempt, due_at, stateOwner: Delivery logsubscription_id, failures, paused_untilOwner: Control planeExternal interfaces
/v1/subscriptionsCreate endpoint, filters, and secret
/internal/eventsPublish one immutable domain event
/v1/deliveries/{id}/replayCreate an authorized new delivery chain
/v1/deliveries?cursor={cursor}Inspect attempts and responses
03Deep dives and trade-offsChoose one
Retry policy
Use exponential backoff with jitter, maximum age, and Retry-After support
Retry forever is a denial-of-service featureOrdering
Offer per-subscription ordered mode only when required and expose head-of-line cost
Global ordering is unnecessaryTenant isolation
Partition quotas by tenant and endpoint, cap concurrent connections, and pause unhealthy targets
One customer must not consume the delivery fleet04Failures, recovery, and evidence3 scenarios
Endpoint times out
Record ambiguous attempt and retry with same event identity
timeout rate and oldest pendingSecret rotation
Sign with versioned active secret and allow bounded overlap
verification failures by secret versionDNS or certificate failure
Back off endpoint-wide and surface clear health state
paused endpoints05What makes the answer seniorInterviewer signals
- At-least-once means the receiver needs event IDs and idempotency
- The operational product includes logs, replay, and endpoint health
- Per-endpoint ordering is a paid complexity, not a default
- Primary trade-off: Per-endpoint order protects consumers but allows one poison event to block the stream.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.