MentorNode
Start free
Platform & InfrastructureHarddesign-metrics-alerting

Design a Metrics & Alerting System

Design the system that tells you the other systems are broken: metric collection from a huge fleet, rule evaluation at interval, and alert routing that doesn't page ten people for one outage.

Pull vs Push CollectionTime-Series StorageRule EvaluationAlert Deduplication
Traffic & Capacity Estimates:

50k hosts · 5M active series · 10k alert rules evaluated every 30s

Functional Requirements

  • •Collect metrics from every service and host with labels, at a configurable interval.
  • •Evaluate alert rules on a schedule and fire when a condition holds for a duration.
  • •Deduplicate, group, and route alerts to the right on-call target with escalation.
  • •Support silences, maintenance windows, and inhibition of downstream alerts.

Non-Functional Requirements

  • •The monitoring system must not fail with the systems it monitors — separate failure domain.
  • •Alert latency under 60 seconds from condition to page.
  • •A metric cardinality spike must degrade that tenant, not the platform.

Back-of-the-Envelope Math

  • 5M series at a 15s scrape interval = ~333k samples/second sustained.
  • 10k rules over 30s windows means the evaluation engine reads millions of samples every cycle.

Key Architectural Trade-offs

  • Pull-based scraping gives you free liveness detection and needs service discovery; push-based works behind NAT and makes 'stopped reporting' ambiguous.
  • A database outage triggers a hundred correlated alerts — inhibition rules and grouping are what turn that into one actionable page.
  • Long retention at full resolution is what you want during a postmortem and is precisely what nobody wants to pay for — hence downsampling tiers.

Click or drag a component onto the canvas, then connect the handles to draw the data flow.

3 nodes · 2 edges

Components · 35

Client & Edge4
Compute & Gateway7
Storage & Caching11
Messaging & Streaming6
Coordination & Ops5
Intelligence2
Canvas overview