04QuestionsWebhook delivery system

04 · Worked prompt

Webhook delivery system

Design subscriptions, signing, durable delivery, exponential retry, idempotency, and tenant isolation.
35 minInterview blueprint
INTERVIEW RUBRIC

What a passing answer must show

100 points · 45 minutes

  1. 20pts

    Scope the problem

    0–5 min

    Prioritize the core flows, state the scale, and name the non-goals.

  2. 15pts

    Define contracts

    5–10 min

    Identify durable entities, APIs, idempotency, and the source of truth.

  3. 30pts

    Complete the diagram

    10–25 min

    Trace one write path and one read path. Label the commit boundary and async work.

  4. 20pts

    Lead one deep dive

    25–38 min

    Choose the highest-risk trade-off and explain the mechanism, alternative, and cost.

  5. 15pts

    Prove reliability

    38–45 min

    Walk a failure, recovery, metric, bottleneck, and evolution path.

DRAW THIS FIRST

One complete box-and-arrow design

Design subscriptions, signing, durable delivery, exponential retry, idempotency, and tenant isolation.

Webhook delivery system · system architecture
Webhook delivery system system architecture. Each delivery attempt is durable; signing, backoff, and endpoint isolation make at-least-once delivery operable. Request path: Event producers to Webhook API to Delivery planner to Delivery DB. Asynchronous path: Due-time queues to Delivery workers. Read path: Delivery status API to Endpoint health view. External dependency: Subscriber endpoints.

Write

Event producers enters through Webhook API. Delivery planner owns validation and commits the durable record to Delivery DB.

Propagate

Due-time queues separates the committed write from background work. Delivery workers can retry safely while it builds Endpoint health view.

Read

Delivery status API serves from Endpoint health view, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Subscriber endpoints as an explicit dependency.

Say this first: Each delivery attempt is durable; signing, backoff, and endpoint isolation make at-least-once delivery operable.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

Each delivery attempt is durable; signing, backoff, and endpoint isolation make at-least-once delivery operable.

INTERVIEW CONTRACT

Design subscriptions, signing, durable delivery, exponential retry, idempotency, and tenant isolation.

CAPACITY QUESTIONS TO QUANTIFY

At-least-once · seconds first attempt · days retry · strict tenant quotas. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Event producers → Webhook API

    Publish tenant events enters over HTTPS / RPC. Webhook API handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Webhook API → Delivery planner → Delivery DB

    Delivery planner receives the command, checks invariants and retry identity, then uses create deliveries to update Delivery DB. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Delivery planner → Due-time queues → Delivery workers → Endpoint health view

    Delivery planner emits publish after commit; Delivery workers uses consume and record outcome to build Endpoint health view. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Webhook API → Delivery status API → Endpoint health view / Delivery DB

    Delivery status API uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    Delivery workers → Subscriber endpoints

    signed POST crosses into Subscriber endpoints. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Webhook APIAuth + subscription versionIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Delivery plannerMatch subscriptions + due timeWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Delivery DBEvent, endpoint, attempt stateAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Due-time queuesTenant + next-attempt bucketDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Delivery workersSign + bounded HTTP sendReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Endpoint health viewLatency, failures, pauseRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Delivery status APIAttempts + terminal reasonRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Subscriber endpointsIdempotent HTTPS receiversExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
PostgreSQL stores subscriptions/deliveries/attempts; Kafka/SQS stores due work; object storage retains large payloads.
Partitioning / sharding
Partition by tenant/endpoint + time bucket; preserve per-endpoint order only if required and isolate huge tenants.
Indexes
Unique (event_id, endpoint_id), due (state, next_attempt, shard), attempts by delivery/time, and endpoint failure streak.
Replication + consistency
Producer acceptance requires durable delivery creation. Delivery is at least once with stable event IDs and conditional local transitions.
Cache, queue + recovery
Queues carry delivery IDs; use visibility leases, exponential backoff+jitter, per-endpoint concurrency, circuit breaking, and DLQs.
Capacity math
Estimate events/sec × subscribers, payload size, endpoint latency, retry amplification, retention, and worst endpoint.
Alternative rejected
Synchronous fan-out couples producer availability to subscribers; durable delivery records isolate slow and failing endpoints.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Retry policy

Use exponential backoff with jitter, maximum age, and Retry-After support

Retry forever is a denial-of-service feature
Ordering

Offer per-subscription ordered mode only when required and expose head-of-line cost

Global ordering is unnecessary
Tenant isolation

Partition quotas by tenant and endpoint, cap concurrent connections, and pause unhealthy targets

One customer must not consume the delivery fleet
05

Failure pressure test

Show detection, containment, recovery, and evidence.

Endpoint times out

Record ambiguous attempt and retry with same event identity

timeout rate and oldest pending
Secret rotation

Sign with versioned active secret and allow bounded overlap

verification failures by secret version
DNS or certificate failure

Back off endpoint-wide and surface clear health state

paused endpoints
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution
SAY THIS WHILE YOU DRAW

A four-part talk track

  1. Scope

    “I’ll prioritize register event subscriptions and deliver signed events at least once.”

  2. Scale

    “The design changes around at-least-once · seconds first attempt · days retry · strict tenant quotas.”

  3. Decision

    “Per-endpoint order protects consumers but allows one poison event to block the stream.”

  4. Risk

    “The first failure I want to pressure-test is: Slow receivers and unbounded retries can consume the entire fleet.”

Reference details

Open these only after you can explain the diagram above without reading.

01Requirements and state lifecycle4 requirements
  • Register event subscriptions
  • Deliver signed events at least once
  • Retry transient failures for days
  • Protect the fleet from slow or broken tenant endpoints
Webhook delivery system · state lifecycle
02Data model and APIs4 entities · 4 interfaces

Core entities

Subscriptionsubscription_id, tenant_id, endpoint, secret_version, filtersOwner: Subscription service
WebhookEventevent_id, type, payload_ref, occurred_atOwner: Event store
Deliverydelivery_id, event_id, subscription_id, attempt, due_at, stateOwner: Delivery log
EndpointHealthsubscription_id, failures, paused_untilOwner: Control plane

External interfaces

POST /v1/subscriptions

Create endpoint, filters, and secret

POST /internal/events

Publish one immutable domain event

POST /v1/deliveries/{id}/replay

Create an authorized new delivery chain

GET /v1/deliveries?cursor={cursor}

Inspect attempts and responses

03Deep dives and trade-offsChoose one

Retry policy

Use exponential backoff with jitter, maximum age, and Retry-After support

Retry forever is a denial-of-service feature

Ordering

Offer per-subscription ordered mode only when required and expose head-of-line cost

Global ordering is unnecessary

Tenant isolation

Partition quotas by tenant and endpoint, cap concurrent connections, and pause unhealthy targets

One customer must not consume the delivery fleet
04Failures, recovery, and evidence3 scenarios

Endpoint times out

Record ambiguous attempt and retry with same event identity

timeout rate and oldest pending

Secret rotation

Sign with versioned active secret and allow bounded overlap

verification failures by secret version

DNS or certificate failure

Back off endpoint-wide and surface clear health state

paused endpoints
05What makes the answer seniorInterviewer signals
  • At-least-once means the receiver needs event IDs and idempotency
  • The operational product includes logs, replay, and endpoint health
  • Per-endpoint ordering is a paid complexity, not a default
  • Primary trade-off: Per-endpoint order protects consumers but allows one poison event to block the stream.
BEFORE THE NEXT QUESTION

Can you redraw it from memory?

  • Name the source of truth.
  • Trace the write and read paths.
  • Defend one trade-off.
  • Recover from one failure.