04QuestionsJob scheduler
04 · Worked prompt
Job scheduler
Design durable schedules, leases, worker heartbeats, retries, priorities, and exactly-once effects.What a passing answer must show
100 points · 45 minutes
- 20pts
Scope the problem
0–5 minPrioritize the core flows, state the scale, and name the non-goals.
- 15pts
Define contracts
5–10 minIdentify durable entities, APIs, idempotency, and the source of truth.
- 30pts
Complete the diagram
10–25 minTrace one write path and one read path. Label the commit boundary and async work.
- 20pts
Lead one deep dive
25–38 minChoose the highest-risk trade-off and explain the mechanism, alternative, and cost.
- 15pts
Prove reliability
38–45 minWalk a failure, recovery, metric, bottleneck, and evolution path.
One complete box-and-arrow design
Design durable schedules, leases, worker heartbeats, retries, priorities, and exactly-once effects.

Write
Job owners enters through Scheduler API. Schedule service owns validation and commits the durable record to Schedule/job DB.
Propagate
Priority ready queues separates the committed write from background work. Worker fleet can retry safely while it builds Attempt history.
Read
Status service serves from Attempt history, then checks authoritative state whenever freshness, policy, or correctness requires it. It also consults Target systems as an explicit dependency.
Say this first: Durable occurrences and fenced leases separate scheduling precision from effect uniqueness.
Open the full whiteboard ↗Explain every boundary before adding more boxes.
Durable occurrences and fenced leases separate scheduling precision from effect uniqueness.
Design durable schedules, leases, worker heartbeats, retries, priorities, and exactly-once effects.
Millions timers · second precision · variable runtimes · tenant priorities. State average and peak load, stored bytes, bandwidth or open connections, and the growth horizon before choosing a partitioning strategy.
End-to-end walkthrough
Trace the architecture in this order.
- 01
Enter and classify the request
Job owners → Scheduler APISchedule, pause, inspect enters over HTTPS / RPC. Scheduler API handles identity, admission, routing, and request context; it deliberately does not own domain truth.
- 02
Validate, then cross the commit boundary
Scheduler API → Schedule service → Schedule/job DBSchedule service receives the command, checks invariants and retry identity, then uses create occurrence to update Schedule/job DB. The user-visible mutation is accepted only after this boundary succeeds.
- 03
Move replayable work off the request path
Schedule service → Priority ready queues → Worker fleet → Attempt historySchedule service emits publish after commit; Worker fleet uses consume and record attempt to build Attempt history. Consumers must tolerate duplicate delivery and stale retries because this path is asynchronous.
- 04
Serve reads from the right authority
Scheduler API → Status service → Attempt history / Schedule/job DBStatus service uses optimized read for the common, read-optimized path and strong read when correctness or repair requires authoritative state. The API must state the freshness promise instead of hiding it.
- 05
Contain the dependency boundary
Worker fleet → Target systemsexecute effect crosses into Target systems. Treat timeouts as ambiguous, use a deadline and idempotent retry or reconciliation, and keep the core state recoverable when the dependency is unavailable.
Ownership ledger
Why each box exists—and what it must defend.
| Component | Owns | Why it exists | Interviewer probe |
|---|---|---|---|
| Scheduler APIAuth + tenant quotas | Identity, admission, routing | Protects the system edge and attaches trusted context before domain work begins. | Timeout budgets, quotas, regional routing |
| Schedule servicePersist next occurrence | Write invariants and retry identity | Serializes or conditionally applies state changes before acknowledging success. | Concurrent writes, deduplication, hot ownership |
| Schedule/job DBOccurrences + lease epochs | Authoritative durable state | Provides the one record used to resolve disputes, recover, and rebuild projections. | Partition key, replication, consistency |
| Priority ready queuesDue jobs by class | Durable asynchronous handoff | Absorbs bursts and lets slow or optional work retry independently of the request. | Ordering key, lag, retention, dead letters |
| Worker fleetClaim, heartbeat, execute | Replayable processing | Runs expensive, fan-out, or side-effecting work with leases and bounded retries. | Idempotency, poison work, autoscaling |
| Attempt historyProgress + terminal result | Rebuildable query state | Shapes data for the dominant reads without weakening the write-side invariant. | Freshness, versioning, rebuild time |
| Status serviceSchedules + attempts | Read composition and freshness policy | Chooses authoritative or derived state and returns a stable client contract. | Fan-out, cache policy, partial results |
| Target systemsIdempotent job effects | External capability, not local truth | Keeps a specialized or third-party concern behind a replaceable contract. | Ambiguous timeout, circuit breaking, fallback |
Physical design
Name the database, shard key, indexes, and guarantees.
- Database + storage
- PostgreSQL/distributed SQL stores schedules, occurrences, leases, and attempts; a durable priority queue stores runnable work.
- Partitioning / sharding
- Hash schedules by tenant/schedule ID and bucket due work by time + shard so pollers do not contend on one global index.
- Indexes
- Due (bucket, next_run, shard, state), unique (schedule_id, occurrence_time), attempts by occurrence/state, and lease expiry.
- Replication + consistency
- Claims are strong and fenced. Execution is at least once; exactly-once effects require an idempotency key at the destination.
- Cache, queue + recovery
- A transactional outbox releases occurrences; visibility leases, heartbeat, jittered retry, dead letters, and reconciliation handle ambiguity.
- Capacity math
- Estimate schedules, timers/sec, precision, top-of-hour burst, duration distribution, and oldest-ready age.
- Alternative rejected
- One cron heap works initially but lacks failover and burst isolation; persisted next occurrences plus sharded pollers evolve safely.
Deep-dive candidates
Pick one risk and explain the mechanism, alternative, and cost.
Timer index
Bucket by due time and shard schedule ownership; maintain a short lookahead queue
A full-table poll is not a scaling planLease fencing
Increment a token on every claim and reject stale completion
Expiration alone cannot stop a zombie workerRecurring policy
Define skip, coalesce, or catch-up behavior for missed runs
Cron syntax does not define business semanticsFailure pressure test
Show detection, containment, recovery, and evidence.
Dispatcher restart
Rescan overlap and rely on unique occurrence identity
duplicate occurrence attemptsWorker partition
Expire lease and fence late completion
stale token rejectionsQueue overload
Apply tenant fairness and delay low priority work
schedule delay by priority- Functional requirements and non-goals
- Peak traffic, storage, bandwidth, and growth
- Entities, APIs, idempotency, and pagination
- Source of truth and consistency promise
- Partition key, replicas, caches, and hot spots
- Retries, backpressure, failover, and reconciliation
- Latency, saturation, correctness, and recovery metrics
- Security, migration, cost, and multi-region evolution
A four-part talk track
- Scope
“I’ll prioritize create one-time and recurring schedules and dispatch near the requested time.”
- Scale
“The design changes around millions timers · second precision · variable runtimes · tenant priorities.”
- Decision
“Indexed polling is simple; delay queues or time wheels scale huge timer populations.”
- Risk
“The first failure I want to pressure-test is: Clock skew and lease races can run one job more than once.”
Reference details
Open these only after you can explain the diagram above without reading.
01Requirements and state lifecycle4 requirements
- Create one-time and recurring schedules
- Dispatch near the requested time
- Support priorities, retries, cancellation, and history
- Recover from dispatcher and worker failure
Each transition must be durable, observable, and safe to retry.
02Data model and APIs4 entities · 3 interfaces
Core entities
schedule_id, rule, next_run_at, policy, versionOwner: Schedule storejob_id, schedule_id, due_at, priority, stateOwner: Dispatcherjob_id, attempt, lease_token, deadlineOwner: Worker systemjob_id, outcome, effect_keyOwner: Result storeExternal interfaces
/v1/schedulesCreate an idempotent schedule
/v1/schedules/{id}Version-check pause, resume, or rule change
/internal/jobs/{id}/completeCommit a result under the current lease token
03Deep dives and trade-offsChoose one
Timer index
Bucket by due time and shard schedule ownership; maintain a short lookahead queue
A full-table poll is not a scaling planLease fencing
Increment a token on every claim and reject stale completion
Expiration alone cannot stop a zombie workerRecurring policy
Define skip, coalesce, or catch-up behavior for missed runs
Cron syntax does not define business semantics04Failures, recovery, and evidence3 scenarios
Dispatcher restart
Rescan overlap and rely on unique occurrence identity
duplicate occurrence attemptsWorker partition
Expire lease and fence late completion
stale token rejectionsQueue overload
Apply tenant fairness and delay low priority work
schedule delay by priority05What makes the answer seniorInterviewer signals
- Exactly-once execution is unrealistic; exactly-once effect is the goal
- Always define missed-run behavior
- The lease token belongs in the completion transaction
- Primary trade-off: Indexed polling is simple; delay queues or time wheels scale huge timer populations.
Can you redraw it from memory?
- Name the source of truth.
- Trace the write and read paths.
- Defend one trade-off.
- Recover from one failure.