Rate Limiting

SYSTEM DESIGN SERIES  ยท  TOPIC 18 OF 23+  ยท  part of the 17 โ†’ 18 โ†’ 19 chain
1. The Same Burst, Two Different Limits
A limit of 3 - but watch what actually gets through right at the edge (looping)
Fixed window: each window checks its own count - 3 fits, so 3 fits, twice
Fixed window: the counter resets hard at the boundary, no memory of what just happened
Fixed window: 6 requests in ~2 seconds - double the "3 per window" you configured
Token bucket: only 3 tokens exist, no matter which side of a clock boundary you're on
Token bucket: the bucket empties, and it refills slowly - not all at once
Token bucket: the burst is capped at 3, full stop, regardless of timing
2. Why Rate Limit At All?
Not every request deserves to get in, right now
  • Protects finite resources - CPU, connections, disk I/O (topic 03)
  • Stops retry storms from becoming self-inflicted outages (topic 17)
  • Fairness - one noisy client shouldn't starve everyone else
  • Cost control - quotas keep usage, and bills, predictable
๐Ÿšช A bouncer doesn't make the club bigger. It just decides who waits outside so the people already in can actually move.
3. Token Bucket vs Leaky Bucket
Token Bucket
  • Tokens refill steadily, up to a max
  • Unused tokens let you burst later
  • Allows bursts, caps the average
VS
Leaky Bucket
  • Requests queue up, drain at a fixed rate
  • Output is smooth no matter the input
  • No bursts at all - ever, by design
๐Ÿชฃ Analogy: a token bucket is a jar of tickets you can save up and spend all at once. A leaky bucket is a funnel - pour in fast or slow, it drips out at the exact same rate either way.
4. The Algorithms at a Glance
Four ways to answer "has this client had too many turns?"
Fixed Window
Dead simple counter per bucket. The boundary can double your real limit.
Sliding Window
Counts the actual trailing window, not a fixed clock tick. No boundary flaw.
Token Bucket
Refills steadily, spends in bursts. The most commonly reached-for default.
Leaky Bucket
Queues and drains at a constant rate. Smoothest output, adds latency.
5. Where It's Enforced & What the Client Sees
Limits live somewhere specific, and they say so
  • Client-side - cooperative, easy to bypass, still worth doing
  • Gateway / load balancer - one place to enforce it for everyone (topic 06)
  • Distributed - many nodes need a shared counter (often Redis) to agree at all
  • The signal: HTTP 429 Too Many Requests, with a Retry-After header
6. What Gets Limited
The same algorithm, applied at different scopes
Per IP Address
Per User / API Key
Per Endpoint
Global / System-Wide
Same rule applied wherever a single bad actor or a single hot path could hurt
Most real systems stack these. A per-IP limit, a per-key limit, and a global ceiling can all be true at the same time.
7. Choosing Your Strategy
Allow bursts when:
โœ” Legitimate traffic is naturally bursty (batch uploads)
โœ” The backend can absorb short spikes fine
โœ” Occasional idle capacity is fine to spend all at once
Rate Limiting Decides Who Waits - Before the System Decides For You
Smooth to constant when:
โœ” The downstream resource is fragile, any burst hurts
โœ” Predictable, steady load matters more than peak throughput
โœ” A little added latency is an acceptable trade
๐Ÿ’ก A rate limiter shared across many servers needs those servers to agree on the count. That agreement problem is Consensus & Leader Election - topic 19, next.