Design an API Rate Limiter, stage 10 of 10: defend it
Defend a 429
A Pro customer files a ticket: "Our logs show we sent 540 requests in the last minute, under our 600 limit, and we got 429s. Your limiter is broken."
Respond as the engineer who built it: how can that happen, which cases are bugs, and what would you check?
System so far· 6 parts
Select a component to see what it is responsible for and which state it owns.
- 1API clients → Load balancer: Requests with API key
- 2Load balancer → API instances: Round-robin across instances
- 3API instances → Redis: Atomic take-tokens script
- 4API instances → Postgres: Admitted requests; plan lookups (cached)
- 5API instances → Search cluster: Admitted searches, cost-weighted
What you need to know
Customers read plans as "600 requests a minute". A token bucket enforces something slightly different: a burst of up to B at once, then a refill of r per second. Over a long period those agree. Over a few seconds they can differ.
Think first
Pro has B = 50 and r = 10 per second. A client sends 200 requests in the first 5 seconds of a minute, then 340 spread over the remaining 55 seconds. Does it get any 429s?Other reasons a client's count and the limiter's count disagree:
- The limiter ran in degraded mode (local limits during a Redis failure).
- The client's HTTP library retried requests that its own logs don't show.
- Other services share the same API key.
- The client counts by send time, the server by arrival time.
Each of these is checkable if the limiter records its decisions: per key, whether a request was allowed, the tokens remaining, and whether degraded mode was on.
Check
Which observation would point to a real bug rather than expected behaviour?