MentorNode
Start free
Messaging & Event-DrivenMediumdesign-job-scheduler

Design a Distributed Job Scheduler

Design cron for a fleet: recurring and one-off jobs that fire close to on time, exactly once, and keep firing when the node that owned them disappears.

Leader ElectionTimer WheelsLeasesAt-Least-Once Execution
Traffic & Capacity Estimates:

50M scheduled jobs · 20k firings/second at peak · ±1 second firing accuracy

Functional Requirements

  • •Register recurring (cron expression) and one-off (run-at) jobs with a payload and a target.
  • •Dispatch due jobs to a worker pool and track execution state, retries, and failures.
  • •Prevent concurrent execution of the same job instance across the fleet.
  • •Support cancellation, pause, and backfill of missed runs after an outage.

Non-Functional Requirements

  • •Firing accuracy within 1 second of the scheduled time at p99.
  • •No job silently skipped — a missed window must be visible and recoverable.
  • •Scheduler node failure must not delay unrelated jobs by more than the failover window.

Back-of-the-Envelope Math

  • 50M jobs partitioned into per-minute time buckets means scanning one small bucket per tick rather than a 50M-row table.
  • 20k firings/s at peak with an average 200ms job means ~4,000 concurrently executing workers.

Key Architectural Trade-offs

  • Database polling by due-time index (simple, easy to reason about, load grows with table size) vs an in-memory timer wheel (O(1) firing, needs durable rebuild on restart).
  • Exactly-once firing is unachievable across failures — leases plus idempotent jobs is the buildable version.
  • Leader-per-partition scales horizontally and needs coordination; a single leader is simple and caps throughput at one node.

Click or drag a component onto the canvas, then connect the handles to draw the data flow.

3 nodes · 2 edges

Components · 35

Client & Edge4
Compute & Gateway7
Storage & Caching11
Messaging & Streaming6
Coordination & Ops5
Intelligence2
Canvas overview