04QuestionsYelp
04 · Worked prompt
Yelp
Design proximity search, geospatial indexing, ranking, reviews, and read-heavy caching.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design proximity search, geospatial indexing, ranking, reviews, and read-heavy caching.

Write
Search + owner clients enters through Local API edge. Business/review service owns validation and commits the durable record to Business/review DB.
Propagate
Change data stream separates the committed write from background work. Index/ranking workers can retry safely while it builds Search/geo indexes.
Read
Discovery service serves from Search/geo indexes, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Media object store as an explicit dependency.
Say this first: Business and review records are truth; text, spatial, and ranking structures are convergent query indexes.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
Business and review records are truth; text, spatial, and ranking structures are convergent query indexes.
Design proximity search, geospatial indexing, ranking, reviews, and read-heavy caching.
Read-heavy · skewed urban density · low-ms candidates · frequent updates. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
Search + owner clients → Local API edgeNear-me query, updates enters over HTTPS / RPC. Local API edge handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Local API edge → Business/review service → Business/review DBBusiness/review service receives the command, checks invariants and retry identity, then uses source transaction to update Business/review DB. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Business/review service → Change data stream → Index/ranking workers → Search/geo indexesBusiness/review service emits publish after commit; Index/ranking workers uses consume and project / update to build Search/geo indexes. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Local API edge → Discovery service → Search/geo indexes / Business/review DBDiscovery service uses retrieve + rank for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
Business/review service → Media object storedependency call crosses into Media object store. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Local API edgeAuth + geo context | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Business/review serviceValidate + commit source record | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Business/review DBCanonical records | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| Change data streamVersioned source mutations | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| Index/ranking workersGeo cells + text + aggregates | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Search/geo indexesBounded candidate retrieval | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Discovery serviceSpatial + text + ranking | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Media object storePhotos + moderation state | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- PostgreSQL owns businesses/reviews; OpenSearch serves text + geo candidates; Redis caches hot result sets; object storage/CDN serves media.
- Partitioning / sharding
- Truth partitions by business/owner; search routes by geography plus hash so local queries touch bounded shards.
- Indexes
- SQL owner/uniqueness, spatial/inverted/facet indexes, and cursors over score + business_id.
- Replication + consistency
- Edits/reviews commit strongly. Search/rating aggregates are eventual, versioned projections rebuilt from CDC.
- Cache, queue + recovery
- CDC compares source versions; cache cell/category/query combinations briefly and coalesce repeated map-pan requests.
- Capacity math
- Estimate businesses, reviews/sec, searches/sec, radius, dense-city documents/cell, index lag, and media egress.
- Alternative rejected
- PostGIS may suffice initially; fuzzy text, facets, geo ranking, and high query scale earn a dedicated search index.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
Spatial retrieval
Use hierarchical cells and expand rings until enough candidates are found
Fixed-radius scans behave poorly across density extremesRanking
Separate eligibility filters from learned or weighted ranking
A closed or unsafe business should not be rescued by scoreReview trust
Keep raw reviews, moderation state, and versioned aggregates
Average rating alone is easy to manipulateFailure pressure test
Show detection, containment, recovery, and evidence.
Dense cell
Use finer child cells and a candidate cap
candidates per cellIndex lag
Verify critical business status during hydration
stale-status filtersBad review burst
Quarantine suspicious reviews and delay aggregate effect
trust-filter rate- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize search nearby businesses and categories and filter by open-now and attributes.”
- Scale
“The design changes around read-heavy · skewed urban density · low-ms candidates · frequent updates.”
- Decision
“Small spatial cells reduce scans but increase boundary fan-out and update churn.”
- Risk
“The first failure I want to pressure-test is: Fake reviews and dense cells can distort or overload results.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Search nearby businesses and categories
- Filter by open-now and attributes
- Rank relevant, trusted results
- Ingest reviews and business changes without blocking reads
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 3 interfaces
Core entities
business_id, location, categories, attributes, versionOwner: Business storereview_id, business_id, rating, trust_stateOwner: Review servicecell_id, business_ids, densityOwner: Spatial indexbusiness_id, aggregates, generated_atOwner: Feature storeExternal interfaces
/v1/discovery?query={q}&near={lat,lng}&cursor={c}Return ranked nearby businesses
/v1/businesses/{id}/reviewsCreate a moderated review
/v1/businesses/{id}Version-check a business update
03Deep dives and trade-offsChoose one
Spatial retrieval
Use hierarchical cells and expand rings until enough candidates are found
Fixed-radius scans behave poorly across density extremesRanking
Separate eligibility filters from learned or weighted ranking
A closed or unsafe business should not be rescued by scoreReview trust
Keep raw reviews, moderation state, and versioned aggregates
Average rating alone is easy to manipulate04Failures, recovery, and evidence3 scenarios
Dense cell
Use finer child cells and a candidate cap
candidates per cellIndex lag
Verify critical business status during hydration
stale-status filtersBad review burst
Quarantine suspicious reviews and delay aggregate effect
trust-filter rate05What makes the answer seniorInterviewer signals
- Candidate generation and ranking are different stages
- Discuss boundary cells and urban density skew
- Freshness of ‘open now’ matters more than freshness of descriptive text
- Primary trade-off: Small spatial cells reduce scans but increase boundary fan-out and update churn.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.