Design an API Rate Limiter, stage 6 of 10: break it
Redis goes down
The limiter now sits on every request. The requirement is explicit: if its state store fails, the API stays up.
System so far· 5 parts
Select a component to see what it is responsible for and which state it owns.
- 1API clients → Load balancer: Requests with API key
- 2Load balancer → API instances: Round-robin across instances
- 3API instances → Redis: Atomic take-tokens script
- 4API instances → Postgres: Admitted requests; plan lookups (cached)
What you need to know
When a dependency on the request path fails, every request has to do something. There are two broad choices:
- Fail closed: treat "I can't check" as "no". Safe for the thing being protected, but the dependency's outage becomes your outage.
- Fail open: treat "I can't check" as "yes". The service stays up, but the protection is gone for as long as the dependency is.
The right choice depends on what the check protects. A payment fraud check might fail closed. A limiter whose job is availability usually shouldn't.
A dependency that is slow is often worse than one that is down. A down Redis returns errors in microseconds. A Redis in the middle of a failover may take seconds to answer, and every request waits that long.
A timeout turns slow into failed. It needs to be shorter than the latency the caller can afford to add, here a few milliseconds.
Work it out
The API handles 40,000 requests a second. Redis stops answering and each limit check waits for a 2-second client timeout. About how many requests are stuck waiting at once?Think first
Redis is unavailable, and each of 30 instances must limit a Pro key (600 a minute) using only its own memory. What limit should each instance apply so the fleet stays near the plan?