MentorNode
Start free
Messaging & Event-DrivenHarddesign-message-queue

Design a Distributed Message Queue (Kafka)

Design the broker itself: a partitioned, replicated, append-only log that guarantees ordering within a partition and lets consumers rewind.

Partitioned LogConsumer GroupsOffset CommitsISR Replication
Traffic & Capacity Estimates:

5M messages/second · 1 PB retained · 10k partitions · 7-day replay window

Functional Requirements

  • •Publish to a topic partition and persist durably before acknowledging the producer.
  • •Distribute partitions across a consumer group and rebalance when members join or leave.
  • •Track per-consumer-group offsets so a restarted consumer resumes where it left off.
  • •Retain messages by time or size and allow replay from any retained offset.

Non-Functional Requirements

  • •Publish p99 under 10ms with replication factor 3 and acks=all.
  • •Strict ordering within a partition; no ordering guarantee is offered across partitions.
  • •A broker loss must not lose acknowledged messages or stall the partitions it led.

Back-of-the-Envelope Math

  • 5M msgs/s * 1 KB = 5 GB/s ingress; at RF=3 that's 15 GB/s of disk write across the cluster.
  • 1 PB retained at RF=3 = 3 PB raw; at 20 TB per broker, ~150 brokers.

Key Architectural Trade-offs

  • Partition count is the parallelism ceiling and the rebalance cost — too few caps throughput, too many blow up metadata and leader election time.
  • acks=all with a minimum in-sync replica set is durable and pays the slowest follower's latency; acks=1 is fast and loses data on leader failure.
  • Log-based brokers (replayable, ordered, consumer-tracked offsets) vs traditional queues with per-message ack and dead-letter semantics — replay against per-message redelivery.

Click or drag a component onto the canvas, then connect the handles to draw the data flow.

3 nodes · 2 edges

Components · 35

Client & Edge4
Compute & Gateway7
Storage & Caching11
Messaging & Streaming6
Coordination & Ops5
Intelligence2
Canvas overview