Rate limits, and who gets turned away
Token bucket, sliding window and the difference between protecting your service and being fair to your users.
A token bucket refills at a steady rate and allows a burst up to its size, which matches how real clients behave, quiet, then a flurry. A fixed window is trivial to implement and lets a client send double the intended rate across a window boundary. A sliding window fixes that and costs more state. Pick from the burst behaviour you want to permit, then decide where the counter lives, because a per-instance counter with eight instances is eight times the limit you wrote down.
Limit on the identity that matters. Per IP punishes offices and shared networks and does nothing to a distributed client. Per API key or per account is usually right for an API; per account plus per endpoint protects the expensive route without throttling the cheap ones. And limit at more than one layer: a cheap edge limit that sheds obvious abuse before it costs you compute, and a precise application limit that enforces the business rule.
The response is part of the design. 429 with Retry-After, and headers stating the limit and when it resets, lets a well-built client back off exactly and lets you tell an angry integrator precisely what happened. A bare 429 with no timing information guarantees an immediate retry, which is how a rate limit turns into a self-inflicted denial of service.
You should now be able to
- Choose a rate limiting algorithm and explain its burst behaviour
- Decide what to limit on and at which layer
- Return a limit response a client can actually cooperate with
Loading…