05Design patternsWeb crawling and data pipelines

05 · Reusable pattern

Web crawling and data pipelines

Design frontier queues, politeness, deduplication, parsing, lineage, and incremental refresh.
8 minConcept guideReference-informed · independently authored
01

Lesson spine

What you need to understand.

A crawler is a polite, deduplicating feedback loop: fetches discover more URLs, and the frontier decides what deserves capacity next.

01

URL frontier

Combine global priority with per-host queues so valuable pages advance without violating host-specific rate limits.

02

Politeness

Honor robots rules, identify the crawler, bound concurrency per host, back off errors, and avoid synchronized recrawls.

03

URL dedupe

Canonicalize and check a Bloom filter or durable seen set before scheduling; preserve aliases when they carry product meaning.

04

Content dedupe

Hash normalized content or use similarity fingerprints so mirrors and near-duplicates do not waste storage and processing.

05

ETL pipeline

Parse, validate, enrich, and load through durable stages with lineage, retry, and dead-letter handling.

06

Distributed ownership

Shard hosts across crawler workers so one owner enforces politeness while shared storage and checkpoints support recovery.

02

Before the boxes

Frame the decision.

Outcome

What must work

Design frontier queues, politeness, deduplication, parsing, lineage, and incremental refresh.

Scale

What changes the design

URLs/s · bandwidth · hosts · change rate · dedupe · refresh SLO

Boundary

What owns the truth

Identify the component that commits authoritative state, then separate synchronous confirmation from derived work.

Non-goal

What stays simple

Do not add global coordination, multi-region writes, or a specialized store until a requirement earns the complexity.

03

Architecture map

Trace ownership, not just traffic.

Web crawling and data pipelines · concept mechanism

Walk one representative request across every arrow. Say whether the handoff is synchronous or asynchronous, what identity makes a retry safe, and which step changes authoritative state.

  1. 01

    Seed service — Adds URLs Define the output contract before moving to the next owner.

  2. 02

    Frontier — Schedules by host Define the output contract before moving to the next owner.

  3. 03

    Fetcher — Respects robots and limits Define the output contract before moving to the next owner.

  4. 04

    Parser — Extracts content and links Define the output contract before moving to the next owner.

  5. 05

    Canonicalizer — Deduplicates Define the output contract before moving to the next owner.

  6. 06

    Content store — Versions pages Define the output contract before moving to the next owner.

  7. 07

    Refresh planner — Chooses recrawl Confirm the result and emit the evidence needed to reconcile it.

04

Decision table

Make the trade-offs explicit.

DecisionDefensible positionCost to acknowledge
Primary mechanismBreadth-first improves coverage; priority scheduling spends capacity on valuable pages.The stronger guarantee usually adds coordination, latency, state, or operational work.
Sync vs. asyncKeep only correctness-critical confirmation synchronous. Move derived views, notifications, analytics, and cleanup behind a durable boundary.Async work needs idempotency, lag monitoring, replay, and a product definition for partial completion.
Simple vs. scaledBegin with one logical owner and a clear API. Partition or replicate only the resource proven to be the first bottleneck.Migration requires stable identities, versioned contracts, backfill, and a rollback path.
05

Failure review

Design the recovery path.

DetectBoundRetry safelyReconcileLearn

Topic-specific risk

Crawler traps, duplicates, and slow hosts waste capacity or harm sites.

Response

Persist enough identity and state to distinguish retry, resume, compensation, and operator repair.

Dependency timeout

A timeout is ambiguous: the remote side may have failed, succeeded, or still be running.

Response

Use deadlines, bounded backoff with jitter, idempotency keys, and a status or reconciliation path.

Overload or skew

Average capacity can look healthy while a tenant, key, partition, region, or expensive request saturates one owner.

Response

Expose queue depth and hot-key share, apply backpressure, isolate tenants, and degrade optional work before correctness.

06

Evidence + level bar

Prove the design can be operated.

Core signals

Health of the promise

Measure user-visible latency or freshness, correctness drift, saturation, retry volume, and time to recover. Alert on the failed promise—not only CPU.

Mid-level

Complete and clear

Finish the happy path, identify the state owner, choose reasonable building blocks, and explain one scale mechanism.

Senior

Trade-offs and failure

Separate read and write paths, define consistency, explain partitioning, and make duplicate or partial failure safe.

Staff+

Evolution and operations

Discuss multi-region boundaries, migration, tenant isolation, capacity, observability, and how the architecture changes over time.

07

Interview language

Open the deep dive with a claim.

“For Web crawling and data pipelines, the decision I want to make explicit is this: Breadth-first improves coverage; priority scheduling spends capacity on valuable pages. I’ll trace the state-changing path first, show where the result becomes durable, then test the design against the highest-risk failure and our target scale.”

08 · Retrieval check

Can you defend it without the page?

  1. For Web crawling and data pipelines, where is the correctness boundary and which failure would you test first?
  2. Which component owns committed truth, and what event or response proves the commit?
  3. Where is the first scaling or coordination bottleneck under the stated envelope?
  4. What happens after an ambiguous timeout or duplicate operation?
  5. Which complexity would you remove at one hundredth of the scale?