04QuestionsAd click aggregator
04 · Worked prompt
Ad click aggregator
Design high-volume event ingestion with deduplication, streaming aggregation, late data, and auditable totals.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design high-volume event ingestion with deduplication, streaming aggregation, late data, and auditable totals.

Write
Ad clients enters through Ingestion edge. Event intake owns validation and commits the durable record to Partitioned event log.
Propagate
Stream partitions separates the committed write from background work. Window processors can retry safely while it builds Aggregate store.
Read
Reporting API serves from Aggregate store, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Cold object store as an explicit dependency.
Say this first: Durable raw events are truth; windowed aggregates are replayable projections.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
Durable raw events are truth; windowed aggregates are replayable projections.
Design high-volume event ingestion with deduplication, streaming aggregation, late data, and auditable totals.
Millions events/s · 1–5 minute freshness · multi-year audit retention. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
Ad clients → Ingestion edgeSigned click events enters over HTTPS / RPC. Ingestion edge handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Ingestion edge → Event intake → Partitioned event logEvent intake receives the command, checks invariants and retry identity, then uses append to update Partitioned event log. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Event intake → Stream partitions → Window processors → Aggregate storeEvent intake emits publish after commit; Window processors uses consume and project / update to build Aggregate store. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Ingestion edge → Reporting API → Aggregate store / Partitioned event logReporting API uses range scan for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
Partitioned event log → Cold object storearchive crosses into Cold object store. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Ingestion edgeAuth, schema, quotas | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Event intakeTimestamp + dedupe key | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Partitioned event logImmutable raw clicks | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| Stream partitionsCampaign + event time | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| Window processorsDedupe, watermark, correct | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Aggregate storeMinute/hour totals | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Reporting APICampaign/time queries | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Cold object storeAudit + replay archive | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- Kafka/Pulsar is the durable ingest log, S3/Parquet is the immutable audit archive, and ClickHouse/Pinot/Druid serves window aggregates.
- Partitioning / sharding
- Partition the stream by campaign_id for ordering; time-partition cold files. Split only proven celebrity campaigns with deterministic subkeys.
- Indexes
- Deduplicate on event_id for the billing horizon; sort aggregates by (campaign_id, window_start) and manifests by event date/hour.
- Replication + consistency
- Broker RF=3 across zones. Transport is at least once; dashboard totals are freshness-bounded while reconciled billing totals are versioned and auditable.
- Cache, queue + recovery
- Watermarks handle late data; checkpointed consumers build projections; malformed events go to a repair queue. Cache only recent reporting windows.
- Capacity math
- Estimate peak clicks/sec, bytes/event, retention years, late-event rate, hottest-campaign share, and analytical query concurrency.
- Alternative rejected
- SQL counters are simpler at low volume, but hot rows, late corrections, and replay requirements earn an immutable log plus projections.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
Event time
Use watermarks and an allowed-lateness policy; reopen or compensate closed windows rather than dropping late clicks
This shows you understand that network time and business time differDeduplication
Store event IDs for the billing horizon or use partition-local compacted state
Exactly-once transport is not the same as exactly-once billingAuditability
Keep raw immutable events and version every correction
Finance needs an explanation path, not only a fast counterFailure pressure test
Show detection, containment, recovery, and evidence.
Broker lag
Apply admission control and scale consumers by partition lag
oldest event age and per-partition lagDuplicate delivery
Make window updates conditional on event identity
dedupe hit rate and duplicate charge countLate event storm
Route corrections through bounded backfill jobs
watermark delay and adjustment volume- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize accept signed click events with a stable event id and show campaign totals within five minutes.”
- Scale
“The design changes around millions events/s · 1–5 minute freshness · multi-year audit retention.”
- Decision
“Separate fast dashboards from exact billing totals and reconcile them.”
- Risk
“The first failure I want to pressure-test is: Late, duplicate, or reordered events can alter closed windows and invoices.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Accept signed click events with a stable event ID
- Show campaign totals within five minutes
- Produce exact, explainable billing totals
- Correct late data without silently rewriting history
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 3 interfaces
Core entities
event_id, ad_id, campaign_id, user_hash, occurred_atOwner: Event logcampaign_id, window_start, count, versionOwner: Stream processorcampaign_id, period, delta, reasonOwner: Reconcilerevent_id, accepted_at, statusOwner: Ingestion serviceExternal interfaces
/v1/clicksAccept one idempotent click and return an ingestion receipt
/v1/campaigns/{id}/metrics?window=1hRead a freshness-labeled dashboard aggregate
/v1/campaigns/{id}/billing/{period}Read the reconciled auditable total
03Deep dives and trade-offsChoose one
Event time
Use watermarks and an allowed-lateness policy; reopen or compensate closed windows rather than dropping late clicks
This shows you understand that network time and business time differDeduplication
Store event IDs for the billing horizon or use partition-local compacted state
Exactly-once transport is not the same as exactly-once billingAuditability
Keep raw immutable events and version every correction
Finance needs an explanation path, not only a fast counter04Failures, recovery, and evidence3 scenarios
Broker lag
Apply admission control and scale consumers by partition lag
oldest event age and per-partition lagDuplicate delivery
Make window updates conditional on event identity
dedupe hit rate and duplicate charge countLate event storm
Route corrections through bounded backfill jobs
watermark delay and adjustment volume05What makes the answer seniorInterviewer signals
- A strong answer separates dashboard freshness from financial finality
- Do not claim exactly once without naming the atomic boundary
- The best deep dive is late data plus reconciliation, not Kafka product trivia
- Primary trade-off: Separate fast dashboards from exact billing totals and reconcile them.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.