Rate limiting
Capping how fast a client may use a resource, to protect capacity, enforce fairness, and stay within the limits of the systems you depend on.
Performance & scale
Learn it
Without limits, one client can consume a shared resource: a buggy script hammering an API, a tenant monopolising workers, or your own system exceeding a provider's quota and getting cut off for everyone.
A rate limiter tracks usage per key (user, API token, IP, tenant, or "us, towards provider X") and rejects or delays requests over the limit.
The common algorithms:
- Token bucket: holds up to B tokens, refills at r per second; each request takes one. Bursts up to B, sustained rate r. The most common choice.
- Leaky bucket: requests drain at a fixed rate; excess queues or drops. Smooths output in front of a strict downstream limit.
- Fixed window: count per minute. Simple, but allows 2× bursts at window boundaries.
- Sliding window: smooths the boundary at more bookkeeping.
Same requests, four limiters. The limit is 5 requests per 10 seconds. Pick a pattern or click the timeline to add requests. Requests (10)Fixed window10 through · most in any 10 s: 10Sliding log5 through · most in any 10 s: 5Sliding window counter6 through · most in any 10 s: 6Token bucket6 through · most in any 10 s: 615.0s- Fixed window:
- Counts per calendar window; the count resets at 10 s.
- Sliding log:
- Keeps every accepted timestamp from the last 10 s.
- Sliding window counter:
- Weights the previous window's count by how much of it still overlaps: an estimate, slightly off either way.
- Token bucket:
- Holds 5 tokens and refills one every 2 s, so a full bucket plus refills can admit a little over 5 in a 10 s span.
Where it runs matters as much as the algorithm. A limiter inside each of 20 API instances enforces 20× the intended limit unless state is shared (Redis with atomic scripts) or the limit is divided among instances. Shared state adds a dependency on every request, so decide what happens when the store is down: fail open or closed.
Check
10 servers each enforce 100 requests a minute per user in their own memory. Requests are spread evenly. What's a user's real limit?Limiting outbound calls matters too: if a provider allows 100 requests a second, your fleet needs a shared limiter, plus queueing for work that can wait, rather than discovering the limit through 429s.
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Per-instance limits
- Each instance enforces the full limit, multiplying the effective limit by the instance count.
- Limiter store outage
- Every request depends on the limiter; failing closed takes the API down.
- Retry-on-429 storms
- Clients retry immediately and keep the limiter saturated.
Instead, consider
- Concurrency limits
- The scarce resource is simultaneous work (connections, workers) rather than rate.
- Quotas over long periods
- Fairness is about daily or monthly consumption, as with billing tiers.
- Queueing
- Excess work can wait rather than be rejected.
In practice
- Token bucket in Redis
- Atomic refill-and-take with a Lua script or INCR with expiry.
- API gateway / proxy limits
- Envoy, NGINX, cloud gateways.
- Client-side limiter for outbound calls
- Keeps your fleet under a provider's quota.
It assumes
- Requests can be attributed to a meaningful key.
- Rejected clients receive a clear signal (429 with Retry-After) and back off.
Explain it in your own words
Where you practise it
- Rate limiting a public API
What the numbers say · Choose the algorithm · Where does the count live? · The limiter that leaks · Write the atomic bucket · Redis goes down · The biggest customer · Credential stuffing on login · Not all requests cost the same · Defend a 429
- A job queue that keeps working when workers fall behind
- Product analytics over billions of events
- Notifications across email, push and in-app
What the numbers say · The email provider is failing · The announcement and the security alert · Write the sender
- A payment workflow that never double-charges
Further reading
Engineers describing it in systems they run.
- How we built rate limiting capable of scaling to millions of domains
Cloudflare · Cloudflare · Post, Jun 2017
Approximating a sliding window with two counters, counting per data centre instead of globally, and measuring how wrong the approximation actually is.
Related concepts
- Backpressure and capacity
When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.
- Retries, backoff and jitter
Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.
- Partitioning
Splitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.
- Load shedding
Rejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.