04QuestionsJob scheduler

04 · Worked prompt

Job scheduler

Design durable schedules, leases, worker heartbeats, retries, priorities, and exactly-once effects.
35 minInterview blueprint
INTERVIEW RUBRIC

What a passing answer must show

100 points · 45 minutes

  1. 20pts

    Scope the problem

    0–5 min

    Prioritize the core flows, state the scale, and name the non-goals.

  2. 15pts

    Define contracts

    5–10 min

    Identify durable entities, APIs, idempotency, and the source of truth.

  3. 30pts

    Complete the diagram

    10–25 min

    Trace one write path and one read path. Label the commit boundary and async work.

  4. 20pts

    Lead one deep dive

    25–38 min

    Choose the highest-risk trade-off and explain the mechanism, alternative, and cost.

  5. 15pts

    Prove reliability

    38–45 min

    Walk a failure, recovery, metric, bottleneck, and evolution path.

DRAW THIS FIRST

One complete box-and-arrow design

Design durable schedules, leases, worker heartbeats, retries, priorities, and exactly-once effects.

Job scheduler · system architecture
Job scheduler system architecture. Durable occurrences and fenced leases separate scheduling precision from effect uniqueness. Request path: Job owners to Scheduler API to Schedule service to Schedule/job DB. Asynchronous path: Priority ready queues to Worker fleet. Read path: Status service to Attempt history. External dependency: Target systems.

Write

Job owners enters through Scheduler API. Schedule service owns validation and commits the durable record to Schedule/job DB.

Propagate

Priority ready queues separates the committed write from background work. Worker fleet can retry safely while it builds Attempt history.

Read

Status service serves from Attempt history, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Target systems as an explicit dependency.

Say this first: Durable occurrences and fenced leases separate scheduling precision from effect uniqueness.

Open the full whiteboard ↗
DEFEND THE DIAGRAM

Explain every boundary before adding more boxes.

Durable occurrences and fenced leases separate scheduling precision from effect uniqueness.

INTERVIEW CONTRACT

Design durable schedules, leases, worker heartbeats, retries, priorities, and exactly-once effects.

CAPACITY QUESTIONS TO QUANTIFY

Millions timers · second precision · variable runtimes · tenant priorities. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.

01

End-to-end walkthrough

Trace the architecture in this order.

  1. 01
    Enter and classify the request
    Job owners → Scheduler API

    Schedule, pause, inspect enters over HTTPS / RPC. Scheduler API handles identity, admission, routing, and request context; it deliberately does not own domain truth.

  2. 02
    Validate, then cross the commit boundary
    Scheduler API → Schedule service → Schedule/job DB

    Schedule service receives the command, checks invariants and retry identity, then uses create occurrence to update Schedule/job DB. The user-visible mutation is accepted only after this boundary succeeds.

  3. 03
    Move replayable work off the request path
    Schedule service → Priority ready queues → Worker fleet → Attempt history

    Schedule service emits publish after commit; Worker fleet uses consume and record attempt to build Attempt history. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.

  4. 04
    Serve reads from the right authority
    Scheduler API → Status service → Attempt history / Schedule/job DB

    Status service uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.

  5. 05
    Contain the dependency boundary
    Worker fleet → Target systems

    execute effect crosses into Target systems. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.

02

Ownership ledger

Why each box exists—and what it must defend.

ComponentOwnsWhy it existsInterviewer probe
Scheduler APIAuth + tenant quotasIdentity, admission, routingProtects the system edge and attaches trusted context before domain work begins.Timeout budgets, quotas, regional routing
Schedule servicePersist next occurrenceWrite invariants and retry identitySerializes or conditionally applies state changes before acknowledging success.Concurrent writes, deduplication, hot ownership
Schedule/job DBOccurrences + lease epochsAuthoritative durable stateProvides the one record used to resolve disputes, recover, and rebuild projections.Partition key, replication, consistency
Priority ready queuesDue jobs by classDurable asynchronous handoffAbsorbs bursts and lets slow or optional work retry independently of the request.Ordering key, lag, retention, dead letters
Worker fleetClaim, heartbeat, executeReplayable processingRuns expensive, fan-out, or side-effecting work with leases and bounded retries.Idempotency, poison work, autoscaling
Attempt historyProgress + terminal resultRebuildable query stateShapes data for the dominant reads without weakening the write-side invariant.Freshness, versioning, rebuild time
Status serviceSchedules + attemptsRead composition and freshness policyChooses authoritative or derived state and returns a stable client contract.Fan-out, cache policy, partial results
Target systemsIdempotent job effectsExternal capability, not local truthKeeps a specialized or third-party concern behind a replaceable contract.Ambiguous timeout, circuit breaking, fallback
03

Physical design

Name the database, shard key, indexes, and guarantees.

Database + storage
PostgreSQL/distributed SQL stores schedules, occurrences, leases, and attempts; a durable priority queue stores runnable work.
Partitioning / sharding
Hash schedules by tenant/schedule ID and bucket due work by time + shard so pollers do not contend on one global index.
Indexes
Due (bucket, next_run, shard, state), unique (schedule_id, occurrence_time), attempts by occurrence/state, and lease expiry.
Replication + consistency
Claims are strong and fenced. Execution is at least once; exactly-once effects require an idempotency key at the destination.
Cache, queue + recovery
A transactional outbox releases occurrences; visibility leases, heartbeat, jittered retry, dead letters, and reconciliation handle ambiguity.
Capacity math
Estimate schedules, timers/sec, precision, top-of-hour burst, duration distribution, and oldest-ready age.
Alternative rejected
One cron heap works initially but lacks failover and burst isolation; persisted next occurrences plus sharded pollers evolve safely.
04

Deep-dive candidates

Pick one risk and explain the mechanism, alternative, and cost.

Timer index

Bucket by due time and shard schedule ownership; maintain a short lookahead queue

A full-table poll is not a scaling plan
Lease fencing

Increment a token on every claim and reject stale completion

Expiration alone cannot stop a zombie worker
Recurring policy

Define skip, coalesce, or catch-up behavior for missed runs

Cron syntax does not define business semantics
05

Failure pressure test

Show detection, containment, recovery, and evidence.

Dispatcher restart

Rescan overlap and rely on unique occurrence identity

duplicate occurrence attempts
Worker partition

Expire lease and fence late completion

stale token rejections
Queue overload

Apply tenant fairness and delay low priority work

schedule delay by priority
Before you finish, explicitly cover
  • Functional requirements and non-goals
  • Peak traffic, storage, bandwidth, and growth
  • Entities, APIs, idempotency, and pagination
  • Source of truth and consistency promise
  • Partition key, replicas, caches, and hot spots
  • Retries, backpressure, failover, and reconciliation
  • Latency, saturation, correctness, and recovery metrics
  • Security, migration, cost, and multi-region evolution
SAY THIS WHILE YOU DRAW

A four-part talk track

  1. Scope

    “I’ll prioritize create one-time and recurring schedules and dispatch near the requested time.”

  2. Scale

    “The design changes around millions timers · second precision · variable runtimes · tenant priorities.”

  3. Decision

    “Indexed polling is simple; delay queues or time wheels scale huge timer populations.”

  4. Risk

    “The first failure I want to pressure-test is: Clock skew and lease races can run one job more than once.”

Reference details

Open these only after you can explain the diagram above without reading.

01Requirements and state lifecycle4 requirements
  • Create one-time and recurring schedules
  • Dispatch near the requested time
  • Support priorities, retries, cancellation, and history
  • Recover from dispatcher and worker failure
Job scheduler · state lifecycle
02Data model and APIs4 entities · 3 interfaces

Core entities

Scheduleschedule_id, rule, next_run_at, policy, versionOwner: Schedule store
Jobjob_id, schedule_id, due_at, priority, stateOwner: Dispatcher
Attemptjob_id, attempt, lease_token, deadlineOwner: Worker system
ExecutionResultjob_id, outcome, effect_keyOwner: Result store

External interfaces

POST /v1/schedules

Create an idempotent schedule

PATCH /v1/schedules/{id}

Version-check pause, resume, or rule change

POST /internal/jobs/{id}/complete

Commit a result under the current lease token

03Deep dives and trade-offsChoose one

Timer index

Bucket by due time and shard schedule ownership; maintain a short lookahead queue

A full-table poll is not a scaling plan

Lease fencing

Increment a token on every claim and reject stale completion

Expiration alone cannot stop a zombie worker

Recurring policy

Define skip, coalesce, or catch-up behavior for missed runs

Cron syntax does not define business semantics
04Failures, recovery, and evidence3 scenarios

Dispatcher restart

Rescan overlap and rely on unique occurrence identity

duplicate occurrence attempts

Worker partition

Expire lease and fence late completion

stale token rejections

Queue overload

Apply tenant fairness and delay low priority work

schedule delay by priority
05What makes the answer seniorInterviewer signals
  • Exactly-once execution is unrealistic; exactly-once effect is the goal
  • Always define missed-run behavior
  • The lease token belongs in the completion transaction
  • Primary trade-off: Indexed polling is simple; delay queues or time wheels scale huge timer populations.
BEFORE THE NEXT QUESTION

Can you redraw it from memory?

  • Name the source of truth.
  • Trace the write and read paths.
  • Defend one trade-off.
  • Recover from one failure.