Stage 1 of 10 · Model
What the numbers say
Thirty instances, 40,000 requests a second, plans measured per minute. Before choosing a design, check which simple ideas the arithmetic already rules out.
What you need to know first
A rate limit is a promise per key: "this API key gets 600 requests a minute". The hard part is that the key's requests don't arrive at one place. The load balancer spreads them round-robin over every instance, so with 30 instances each one sees about 1/30th of any key's traffic.
A Pro key sends 600 requests a minute, spread evenly over 30 instances. About how many of them does one instance see per minute?
About 20 per minute.
600 ÷ 30 = 20 a minute per instance.
So an instance that counts only its own traffic sees 20 requests from a key that is already at its limit. It has no way to tell, from what it sees, that the key is at 600 fleet-wide.
The simplest limiter counts requests per key in fixed calendar windows: "requests in 12:00:00–12:00:59". At the start of each minute the counter resets to zero.
That counts an average over the window. It says nothing about how the requests are spread inside the window, or across the boundary between two windows.
Limit: 60 per calendar minute. A client sends 60 requests at 12:00:59 and 60 more at 12:01:00. What does a fixed-window counter do?
It admits all 120. The first 60 land in the 12:00 window, which still had room; the next 60 land in a fresh 12:01 window that just reset to zero.
120 requests in about one second, from a client "limited" to 60 a minute.
Two numbers to keep in mind for anything on the request path:
| Operation | Rough cost |
|---|---|
| Redis round trip in the same region | ~0.3 ms |
| Simple Redis operations, one shard | ~100,000 per second |
| This API's latency budget for the limiter | 2 ms at p99 |
A check that touches shared state on every request is affordable here. The question is what happens to every request when that shared state is slow or gone.
Every request now makes one Redis call. Which concern is the real one?
Redis is now on the path of every request, so its latency and failures become the API's.
Before, a Redis problem was a Redis problem. Now every request waits for it. That is the design question the rest of this investigation keeps returning to.
What the stage asks
Which statements hold?
- Holds
If each of 30 instances enforces 600 requests/minute per key in its own memory, a Pro customer can make up to 18,000 requests a minute.
The load balancer spreads the customer's requests over every instance, and each instance gives them a full allowance. The effective limit scales with the fleet, and autoscaling makes it move.
- Holds
A fixed one-minute window with a limit of 60 can admit 120 requests within two seconds.
60 in the last second of one window, 60 in the first second of the next. Fixed windows enforce an average over the window, not a maximum over any interval.
- Depends
Checking every request against Redis means about 40,000 Redis operations a second, which is a problem.
A single Redis shard handles on the order of 100,000 simple operations a second, and a cluster spreads keys over shards, so throughput is fine. What matters is that Redis is now on the critical path of every request: its latency and availability become the API's. That is the real design question.
- Fails
Limiting by client IP is sufficient for the authenticated API.
One corporate NAT can hide hundreds of honest developers behind one IP, while one customer can spread a script across many machines. The plan belongs to the API key, so the limit must too. IP limits are for traffic with no better identity, such as login.
- Depends
The limiter must be exact: admitting the 601st request in a minute is a bug.
Plans are commercial promises, and protecting the database is an engineering goal. Neither is harmed by a few percent of imprecision. Insisting on exactness costs coordination and latency on every request; a limiter that is approximately right and always fast is usually the better product.
The reasoning
- Per-instance limits multiply by the number of instances, so a fleet-wide limit needs shared state.
- Fixed windows enforce an average per window and allow double bursts at the boundary.
- Shared state on every request puts its latency and availability on the API's critical path.
Two conclusions shape everything after this:
- The state must be shared. Per-instance limits multiply by the fleet size, so a key's counter has to live in one place every instance consults, or be split deliberately.
- The shared state is now on the hot path. Every request waits for it, so its latency budget is the 2 ms requirement and its failure mode is the API's failure mode, unless you design otherwise.
And one framing: rate limiting is a Backpressure and capacity mechanism. It turns "more load than the database can take" into "a clear signal to the client that sent it".