Skip to content

Rate limiting

Capping how fast a client may use a resource, to protect capacity, enforce fairness, and stay within the limits of the systems you depend on.

Performance & scale

Learn it

0 of 1 checks done
  1. Without limits, one client can consume a shared resource: a buggy script hammering an API, a tenant monopolising workers, or your own system exceeding a provider's quota and getting cut off for everyone.

    A rate limiter tracks usage per key (user, API token, IP, tenant, or "us, towards provider X") and rejects or delays requests over the limit.

  2. The common algorithms:

    • Token bucket: holds up to B tokens, refills at r per second; each request takes one. Bursts up to B, sustained rate r. The most common choice.
    • Leaky bucket: requests drain at a fixed rate; excess queues or drops. Smooths output in front of a strict downstream limit.
    • Fixed window: count per minute. Simple, but allows 2× bursts at window boundaries.
    • Sliding window: smooths the boundary at more bookkeeping.
  3. Same requests, four limiters. The limit is 5 requests per 10 seconds. Pick a pattern or click the timeline to add requests.
    Requests (10)
    Fixed window10 through · most in any 10 s: 10
    Sliding log5 through · most in any 10 s: 5
    Sliding window counter6 through · most in any 10 s: 6
    Token bucket6 through · most in any 10 s: 6
    15.0s
    Fixed window:
    Counts per calendar window; the count resets at 10 s.
    Sliding log:
    Keeps every accepted timestamp from the last 10 s.
    Sliding window counter:
    Weights the previous window's count by how much of it still overlaps: an estimate, slightly off either way.
    Token bucket:
    Holds 5 tokens and refills one every 2 s, so a full bucket plus refills can admit a little over 5 in a 10 s span.
  4. Where it runs matters as much as the algorithm. A limiter inside each of 20 API instances enforces 20× the intended limit unless state is shared (Redis with atomic scripts) or the limit is divided among instances. Shared state adds a dependency on every request, so decide what happens when the store is down: fail open or closed.

  5. Check

    10 servers each enforce 100 requests a minute per user in their own memory. Requests are spread evenly. What's a user's real limit?

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Per-instance limits
Each instance enforces the full limit, multiplying the effective limit by the instance count.
Limiter store outage
Every request depends on the limiter; failing closed takes the API down.
Retry-on-429 storms
Clients retry immediately and keep the limiter saturated.

Instead, consider

Concurrency limits
The scarce resource is simultaneous work (connections, workers) rather than rate.
Quotas over long periods
Fairness is about daily or monthly consumption, as with billing tiers.
Queueing
Excess work can wait rather than be rejected.

In practice

Token bucket in Redis
Atomic refill-and-take with a Lua script or INCR with expiry.
API gateway / proxy limits
Envoy, NGINX, cloud gateways.
Client-side limiter for outbound calls
Keeps your fleet under a provider's quota.

It assumes

  • Requests can be attributed to a meaningful key.
  • Rejected clients receive a clear signal (429 with Retry-After) and back off.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Backpressure and capacity

    When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.

  • Retries, backoff and jitter

    Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.

  • Partitioning

    Splitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.

  • Load shedding

    Rejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.