MentorNode
Start free
AI / ML InfrastructureMediumdesign-llm-gateway

Design an LLM API Gateway with Quotas

Design the in-house front door to several model providers: token-based quotas, per-tenant budgets, caching, failover, and streaming responses that make cancellation meaningful.

Token AccountingProvider FailoverSemantic CachingStreaming Responses
Traffic & Capacity Estimates:

20k concurrent streams · 5 providers · per-tenant token budgets · sub-second first token

Functional Requirements

  • •Route requests to a provider and model by policy, cost, and availability, with automatic failover.
  • •Meter input and output tokens per tenant and enforce hard and soft budget limits.
  • •Stream responses token-by-token and stop billing immediately on client cancellation.
  • •Cache deterministic and semantically similar requests to avoid paying twice for the same answer.

Non-Functional Requirements

  • •Time to first token under 1 second at p95.
  • •A provider outage must fail over without dropping in-flight streams where possible.
  • •Token accounting must be accurate enough to bill on and must never block the response path.

Back-of-the-Envelope Math

  • 20k concurrent streams held open for 30s average = long-lived connections, so the gateway must be non-blocking throughout.
  • A 30% semantic cache hit rate on a large workload is a direct 30% reduction in provider spend.

Key Architectural Trade-offs

  • Hard quota enforcement before the call protects the budget and rejects users mid-conversation; soft accounting after the fact is friendlier and lets a tenant overspend.
  • Failover to a different provider preserves availability and changes the model's behaviour underneath the user — sometimes acceptable, sometimes a correctness bug.
  • Semantic caching saves real money and can return a subtly wrong answer for a near-miss query; the similarity threshold is a product risk decision, not a tuning knob.

Click or drag a component onto the canvas, then connect the handles to draw the data flow.

3 nodes · 2 edges

Components · 35

Client & Edge4
Compute & Gateway7
Storage & Caching11
Messaging & Streaming6
Coordination & Ops5
Intelligence2
Canvas overview