Skip to content

Design an API Rate Limiter, stage 3 of 10: decide

Where does the count live?

Thirty instances, autoscaling between 20 and 60, round-robin load balancing. Every instance must agree, closely enough, on how many tokens each key has left.

System so far· 4 parts
123CLIENTAPI clientsEDGELoad balancerSERVICEAPI instancesDATABASEPostgres

Select a component to see what it is responsible for and which state it owns.

  1. 1API clients → Load balancer: Requests with API key
  2. 2Load balancer → API instances: Round-robin across instances
  3. 3API instances → Postgres: Admitted requests; plan lookups (cached)

What you need to know

0 of 2 checks done
  1. Wherever the bucket lives, three properties matter:

    • Shared: every instance must consult the same bucket for a key, or the limit multiplies by the number of instances.
    • Fast: the check runs on every request, inside a 2 ms budget.
    • Atomic: many instances update the same key at the same moment, so the read-refill-take-write sequence must not interleave.

    Durability matters much less. If bucket state is lost, buckets start full, which briefly admits a little extra traffic.

  2. Check

    The store holding every token bucket loses all its data. What happens?