Load shedding
Rejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.
Performance & scale
Learn it
An overloaded server that accepts everything does everything slowly. Requests time out after consuming CPU, retries pile on, and goodput (useful work finished in time) falls toward zero while the server is 100% busy. Overload that isn't shed becomes an outage.
Check
At 150% of capacity, which outcome is better for users?Decide early and cheaply:
- Detect overload from a leading signal: concurrency in flight, queue age, worker utilisation. CPU alone is often too late.
- Reject before expensive work: at the edge, before parsing big bodies or taking locks.
- Prioritise: shed low-priority work (analytics, batch) first; reserve capacity for what matters.
- Tell callers what to do: 503 or 429 with
Retry-After. - Drop work nobody is waiting for: a queued request whose client already timed out should be discarded.
Load shedding complements Rate limiting: rate limits enforce fairness per client under normal load; shedding protects the whole system when total demand exceeds capacity, whoever caused it.
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Shedding too late
- The signal lags, so the server is already thrashing before it starts rejecting.
- Expensive rejection
- Requests are fully parsed or authenticated before being rejected, so shedding barely saves anything.
- Retry storms
- Rejected clients retry immediately and multiply the load.
- Shedding the wrong work
- Without priorities, critical requests are dropped as readily as background work.
Instead, consider
- Autoscaling
- Load rises slowly enough for new capacity to arrive before queues explode.
- Queueing
- The burst is short and the work can wait.
- Degrading responses
- A cheaper version of the response (cached, partial) is better than none.
In practice
- Concurrency limiters
- Cap requests in flight per endpoint or per client.
- Priority load shedders
- Reserve a share of workers for critical traffic and reject the rest first.
- Queue deadlines
- Discard queued work older than the caller's timeout.
It assumes
- Work can be ranked, so some of it is safe to refuse.
- Callers handle rejection by backing off rather than retrying at once.
- The overload signal is measured close to where the shedding happens.
Explain it in your own words
Where you practise it
Related concepts
- Backpressure and capacity
When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.
- Rate limiting
Capping how fast a client may use a resource, to protect capacity, enforce fairness, and stay within the limits of the systems you depend on.
- Retries, backoff and jitter
Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.
- Timeouts and unknown outcomes
A timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.