Learning on Web Dev Open is free for all.

System Design & Performance > Scale, consistency and failureRate limits, and who gets turned away
Phase 06Scale, consistency and failure315 of 434

Rate limits, and who gets turned away

Token bucket, sliding window and the difference between protecting your service and being fair to your users.

Concept14 minAI pair

A token bucket refills at a steady rate and allows a burst up to its size, which matches how real clients behave, quiet, then a flurry. A fixed window is trivial to implement and lets a client send double the intended rate across a window boundary. A sliding window fixes that and costs more state. Pick from the burst behaviour you want to permit, then decide where the counter lives, because a per-instance counter with eight instances is eight times the limit you wrote down.

Limit on the identity that matters. Per IP punishes offices and shared networks and does nothing to a distributed client. Per API key or per account is usually right for an API; per account plus per endpoint protects the expensive route without throttling the cheap ones. And limit at more than one layer: a cheap edge limit that sheds obvious abuse before it costs you compute, and a precise application limit that enforces the business rule.

The response is part of the design. 429 with Retry-After, and headers stating the limit and when it resets, lets a well-built client back off exactly and lets you tell an angry integrator precisely what happened. A bare 429 with no timing information guarantees an immediate retry, which is how a rate limit turns into a self-inflicted denial of service.

You should now be able to

  • Choose a rate limiting algorithm and explain its burst behaviour
  • Decide what to limit on and at which layer
  • Return a limit response a client can actually cooperate with
Ask the community

Loading…