Skip to content

Design an API Rate Limiter/Playground

Rate limiting a public API

Click a part to change it, put something in front of it, or kill it. Turn the traffic up. Every number moves as you go; nothing is graded. About 40,000 requests a second at peak across 20–60 API instances; the average is assumed.

Running
Traffic
CLIENTAPI clientsEDGELoad balancerSERVICEAPI instancesCACHERedisDATABASEPostgresSERVICESearch cluster
Requests failing
0%
Backlog growing
none
Instances running
34

What goes through it

OperationOfferedOutcomeWaitUp
API requests40k/sAll served47 ms99.990%
Limiter checks40k/sAll served2 ms99.999%
Database queries7,500/sAll served16 ms99.95%
Search queries2,500/sAll served40 ms99.999%

Wait is the expected time a caller waits; Up is the share of time every part it waits on is running.

What your changes did

Nothing yet. This is the reference design: change something to see what it buys and costs.

At this traffic

Ask AI what your design does

AI

The numbers above come from a simple model. The AI reads your design, what you changed and these numbers, and explains where it holds and where it fails, citing how the engineers who built it did it. It can be wrong.