MentorNode
Start free
Caching & Content DeliveryMediumdesign-distributed-cache

Design a Distributed Cache (Memcached / Redis Cluster)

Design the in-memory tier itself: how keys are partitioned across nodes, what gets evicted, and what happens to correctness when a node disappears mid-traffic.

Consistent HashingLRU EvictionWrite-Through vs Write-BackReplication
Traffic & Capacity Estimates:

10 TB of cached data · 5M ops/second · p99 under 1ms · 200 cache nodes

Functional Requirements

  • •Get, set, and delete keys with a TTL, partitioned across a cluster of memory nodes.
  • •Evict under memory pressure by a configurable policy (LRU, LFU, TTL-first).
  • •Optionally replicate hot partitions so a node loss doesn't cost the whole slice.
  • •Expose per-key and per-node statistics — hit ratio, eviction rate, memory fragmentation.

Non-Functional Requirements

  • •p99 under 1ms for a single-key read, including network.
  • •A node failure must degrade the hit ratio, never return a wrong value.
  • •Adding capacity must not force a global rehash that empties the cache.

Back-of-the-Envelope Math

  • 10 TB across 200 nodes = 50 GB per node; at 1 KB average value, ~10 billion keys.
  • 5M ops/s across 200 nodes = 25k ops/s per node — well within a single-threaded Redis instance.

Key Architectural Trade-offs

  • Cache-aside (application controls correctness, cold misses hit the DB) vs read-through/write-through (uniform, couples the cache into the write path).
  • Replicating hot partitions costs memory and adds a staleness window; not replicating means a node loss lands directly on the database.
  • Eviction by LRU is cheap and usually right; LFU protects genuinely hot keys but pays for frequency tracking on every access.

Click or drag a component onto the canvas, then connect the handles to draw the data flow.

3 nodes · 2 edges

Components · 35

Client & Edge4
Compute & Gateway7
Storage & Caching11
Messaging & Streaming6
Coordination & Ops5
Intelligence2
Canvas overview