Design an API Rate Limiter, stage 3 of 10: decide
Where does the count live?
Thirty instances, autoscaling between 20 and 60, round-robin load balancing. Every instance must agree, closely enough, on how many tokens each key has left.
System so far· 4 parts
Select a component to see what it is responsible for and which state it owns.
- 1API clients → Load balancer: Requests with API key
- 2Load balancer → API instances: Round-robin across instances
- 3API instances → Postgres: Admitted requests; plan lookups (cached)
What you need to know
0 of 2 checks done
Wherever the bucket lives, three properties matter:
- Shared: every instance must consult the same bucket for a key, or the limit multiplies by the number of instances.
- Fast: the check runs on every request, inside a 2 ms budget.
- Atomic: many instances update the same key at the same moment, so the read-refill-take-write sequence must not interleave.
Durability matters much less. If bucket state is lost, buckets start full, which briefly admits a little extra traffic.
Check
The store holding every token bucket loses all its data. What happens?Consider where else the count could live:
- Each instance's memory costs nothing, but it isn't shared.
- Sticky routing, sending every request for a key to one instance, makes local memory shared again. It needs a load balancer that can route by key, and every scale event moves keys between instances.
- The primary database is shared and durable, but every request would write to the very database the limiter exists to protect.
- Redis is shared, about 0.3 ms away, and can run a small script atomically.
Think first
If the bucket were a Postgres row updated in a transaction, what happens when one key receives 500 requests a second from 30 instances?