04QuestionsAd click aggregator

04 · Worked prompt

Ad click aggregator

Design high-volume event ingestion with deduplication, streaming aggregation, late data, and auditable totals.
35 minInterview blueprint
INTERVIEW RUBRIC

What a passing answer must show

100 points · 45 minutes

  1. 20pts

    Scope the problem

    0–5 min

    Prioritize the core flows, state the scale, and name the non-goals.

  2. 15pts

    Define contracts

    5–10 min

    Identify durable entities, APIs, idempotency, and the source of truth.

  3. 30pts

    Complete the diagram

    10–25 min

    Trace one write path and one read path. Label the commit boundary and async work.

  4. 20pts

    Lead one deep dive

    25–38 min

    Choose the highest-risk trade-off and explain the mechanism, alternative, and cost.

  5. 15pts

    Prove reliability

    38–45 min

    Walk a failure, recovery, metric, bottleneck, and evolution path.

DRAW THIS FIRST

One complete box-and-arrow design

Design high-volume event ingestion with deduplication, streaming aggregation, late data, and auditable totals.

Ad click aggregator · system architecture
Ad click aggregator system architecture. Durable raw events are truth; windowed aggregates are replayable projections. Request path: Ad clients to Ingestion edge to Event intake to Partitioned event log. Asynchronous path: Stream partitions to Window processors. Read path: Reporting API to Aggregate store. External dependency: Cold object store.

Write

Ad clients enters through Ingestion edge. Event intake owns validation and commits the durable record to Partitioned event log.

Propagate

Stream partitions separates the committed write from background work. Window processors can retry safely while it builds Aggregate store.

Read

Reporting API serves from Aggregate store, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Cold object store as an explicit dependency.

Say this first: Durable raw events are truth; windowed aggregates are replayable projections.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

Durable raw events are truth; windowed aggregates are replayable projections.

INTERVIEW CONTRACT

Design high-volume event ingestion with deduplication, streaming aggregation, late data, and auditable totals.

CAPACITY QUESTIONS TO QUANTIFY

Millions events/s · 1–5 minute freshness · multi-year audit retention. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Ad clients → Ingestion edge

    Signed click events enters over HTTPS / RPC. Ingestion edge handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Ingestion edge → Event intake → Partitioned event log

    Event intake receives the command, checks invariants and retry identity, then uses append to update Partitioned event log. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Event intake → Stream partitions → Window processors → Aggregate store

    Event intake emits publish after commit; Window processors uses consume and project / update to build Aggregate store. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Ingestion edge → Reporting API → Aggregate store / Partitioned event log

    Reporting API uses range scan for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    Partitioned event log → Cold object store

    archive crosses into Cold object store. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Ingestion edgeAuth, schema, quotasIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Event intakeTimestamp + dedupe keyWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Partitioned event logImmutable raw clicksAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Stream partitionsCampaign + event timeDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Window processorsDedupe, watermark, correctReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Aggregate storeMinute/hour totalsRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Reporting APICampaign/time queriesRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Cold object storeAudit + replay archiveExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
Kafka/Pulsar is the durable ingest log, S3/Parquet is the immutable audit archive, and ClickHouse/Pinot/Druid serves window aggregates.
Partitioning / sharding
Partition the stream by campaign_id for ordering; time-partition cold files. Split only proven celebrity campaigns with deterministic subkeys.
Indexes
Deduplicate on event_id for the billing horizon; sort aggregates by (campaign_id, window_start) and manifests by event date/hour.
Replication + consistency
Broker RF=3 across zones. Transport is at least once; dashboard totals are freshness-bounded while reconciled billing totals are versioned and auditable.
Cache, queue + recovery
Watermarks handle late data; checkpointed consumers build projections; malformed events go to a repair queue. Cache only recent reporting windows.
Capacity math
Estimate peak clicks/sec, bytes/event, retention years, late-event rate, hottest-campaign share, and analytical query concurrency.
Alternative rejected
SQL counters are simpler at low volume, but hot rows, late corrections, and replay requirements earn an immutable log plus projections.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Event time

Use watermarks and an allowed-lateness policy; reopen or compensate closed windows rather than dropping late clicks

This shows you understand that network time and business time differ
Deduplication

Store event IDs for the billing horizon or use partition-local compacted state

Exactly-once transport is not the same as exactly-once billing
Auditability

Keep raw immutable events and version every correction

Finance needs an explanation path, not only a fast counter
05

Failure pressure test

Show detection, containment, recovery, and evidence.

Broker lag

Apply admission control and scale consumers by partition lag

oldest event age and per-partition lag
Duplicate delivery

Make window updates conditional on event identity

dedupe hit rate and duplicate charge count
Late event storm

Route corrections through bounded backfill jobs

watermark delay and adjustment volume
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution
SAY THIS WHILE YOU DRAW

A four-part talk track

  1. Scope

    “I’ll prioritize accept signed click events with a stable event id and show campaign totals within five minutes.”

  2. Scale

    “The design changes around millions events/s · 1–5 minute freshness · multi-year audit retention.”

  3. Decision

    “Separate fast dashboards from exact billing totals and reconcile them.”

  4. Risk

    “The first failure I want to pressure-test is: Late, duplicate, or reordered events can alter closed windows and invoices.”

Reference details

Open these only after you can explain the diagram above without reading.

01Requirements and state lifecycle4 requirements
  • Accept signed click events with a stable event ID
  • Show campaign totals within five minutes
  • Produce exact, explainable billing totals
  • Correct late data without silently rewriting history
Ad click aggregator · state lifecycle
02Data model and APIs4 entities · 3 interfaces

Core entities

ClickEventevent_id, ad_id, campaign_id, user_hash, occurred_atOwner: Event log
CampaignWindowcampaign_id, window_start, count, versionOwner: Stream processor
BillingAdjustmentcampaign_id, period, delta, reasonOwner: Reconciler
IngestionReceiptevent_id, accepted_at, statusOwner: Ingestion service

External interfaces

POST /v1/clicks

Accept one idempotent click and return an ingestion receipt

GET /v1/campaigns/{id}/metrics?window=1h

Read a freshness-labeled dashboard aggregate

GET /v1/campaigns/{id}/billing/{period}

Read the reconciled auditable total

03Deep dives and trade-offsChoose one

Event time

Use watermarks and an allowed-lateness policy; reopen or compensate closed windows rather than dropping late clicks

This shows you understand that network time and business time differ

Deduplication

Store event IDs for the billing horizon or use partition-local compacted state

Exactly-once transport is not the same as exactly-once billing

Auditability

Keep raw immutable events and version every correction

Finance needs an explanation path, not only a fast counter
04Failures, recovery, and evidence3 scenarios

Broker lag

Apply admission control and scale consumers by partition lag

oldest event age and per-partition lag

Duplicate delivery

Make window updates conditional on event identity

dedupe hit rate and duplicate charge count

Late event storm

Route corrections through bounded backfill jobs

watermark delay and adjustment volume
05What makes the answer seniorInterviewer signals
  • A strong answer separates dashboard freshness from financial finality
  • Do not claim exactly once without naming the atomic boundary
  • The best deep dive is late data plus reconciliation, not Kafka product trivia
  • Primary trade-off: Separate fast dashboards from exact billing totals and reconcile them.
BEFORE THE NEXT QUESTION

Can you redraw it from memory?

  • Name the source of truth.
  • Trace the write and read paths.
  • Defend one trade-off.
  • Recover from one failure.