Design an API Rate Limiter, stage 1 of 10: model
What the numbers say
Thirty instances, 40,000 requests a second, plans measured per minute. Before choosing a design, check which simple ideas the arithmetic already rules out.
System so far· 4 parts
Select a component to see what it is responsible for and which state it owns.
- 1API clients → Load balancer: Requests with API key
- 2Load balancer → API instances: Round-robin across instances
- 3API instances → Postgres: Admitted requests; plan lookups (cached)
What you need to know
A rate limit is a promise per key: "this API key gets 600 requests a minute". The hard part is that the key's requests don't arrive at one place. The load balancer spreads them round-robin over every instance, so with 30 instances each one sees about 1/30th of any key's traffic.
Work it out
A Pro key sends 600 requests a minute, spread evenly over 30 instances. About how many of them does one instance see per minute?The simplest limiter counts requests per key in fixed calendar windows: "requests in 12:00:00–12:00:59". At the start of each minute the counter resets to zero.
That counts an average over the window. It says nothing about how the requests are spread inside the window, or across the boundary between two windows.
Think first
Limit: 60 per calendar minute. A client sends 60 requests at 12:00:59 and 60 more at 12:01:00. What does a fixed-window counter do?Two numbers to keep in mind for anything on the request path:
Operation Rough cost Redis round trip in the same region ~0.3 ms Simple Redis operations, one shard ~100,000 per second This API's latency budget for the limiter 2 ms at p99 A check that touches shared state on every request is affordable here. The question is what happens to every request when that shared state is slow or gone.
Check
Every request now makes one Redis call. Which concern is the real one?