MentorNode
Start free
AI / ML InfrastructureHarddesign-ml-inference-platform

Design an ML Model Serving Platform

Design the serving tier for models on expensive accelerators: request batching, version rollout, autoscaling that accounts for cold starts, and a fallback when the GPU pool is saturated.

Dynamic BatchingGPU AutoscalingModel VersioningShadow Deployment
Traffic & Capacity Estimates:

50k inference requests/second · 200 model versions · p99 under 300ms · GPU-bound

Functional Requirements

  • •Serve multiple model versions concurrently and route traffic by version, tenant, or experiment.
  • •Batch incoming requests dynamically to keep accelerator utilization high without blowing the latency budget.
  • •Load and unload models on demand across a heterogeneous GPU fleet.
  • •Shadow new versions against live traffic and compare outputs before promotion.

Non-Functional Requirements

  • •p99 under 300ms including queueing, batching, and inference.
  • •GPU utilization above 60% — idle accelerators are the dominant cost.
  • •Saturation must shed load or fall back to a smaller model, never queue unboundedly.

Back-of-the-Envelope Math

  • A batch size of 32 can raise throughput 10x over single-request inference and adds up to one batch-window of latency.
  • Cold-loading a 30 GB model takes tens of seconds — autoscaling must lead demand, not follow it.

Key Architectural Trade-offs

  • Dynamic batching is the main lever for GPU efficiency and directly spends latency budget — the batch window is the dial between cost and p99.
  • Model-per-replica is simple and wastes memory across many small models; multi-model serving packs the GPU and introduces noisy-neighbour contention.
  • Autoscaling on queue depth reacts to real pressure and is too slow for multi-minute model loads, so warm pools become a cost you pay for latency.

Click or drag a component onto the canvas, then connect the handles to draw the data flow.

3 nodes · 2 edges

Components · 35

Client & Edge4
Compute & Gateway7
Storage & Caching11
Messaging & Streaming6
Coordination & Ops5
Intelligence2
Canvas overview