Retries, backoff and jitter
Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.
Reliability
Learn it
Many failures are transient: a dropped connection, a 503 during a deploy, a lock timeout. Retrying fixes them. Retrying badly creates new failures: synchronized waves, extra load on a struggling dependency, duplicate side effects.
A sound retry policy answers four questions.
- Is it retryable? Transient errors (timeouts, 503, resets) yes. Permanent ones (400, validation, "card declined") no. Unknown outcomes only if the operation is idempotent.
- How long to wait? Exponential backoff (100 ms, 200 ms, 400 ms… up to a cap) gives the dependency room to recover.
- With jitter. If a thousand clients all wait exactly 400 ms, they all return together. Randomise the delay, e.g. "full jitter": a random value between 0 and the backoff.
- How many in total? Cap attempts per request, and ideally set a retry budget (retries at most ~10% of requests).
Work it out
Three layers each make up to 3 attempts at the layer below. One user action can cause up to how many calls to the bottom service?Check
A payment provider responds 'card declined'. Should you retry with backoff?Pair retries with Timeouts and unknown outcomes, and under sustained failure use a circuit breaker that stops calling the dependency for a while.
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Retry storm
- Synchronized retries without jitter hit a recovering service in waves and knock it over again.
- Amplification across layers
- Nested retries multiply load on the deepest dependency.
- Retrying permanent errors
- Wastes capacity and delays the user's real feedback.
- Retrying non-idempotent calls
- Duplicates charges, emails or records.
Instead, consider
- Fail fast and surface the error
- A human or upstream caller is better placed to decide whether to retry.
- Queue for later processing
- The work can wait; a durable queue retries without holding a request open.
- Hedged requests
- Tail latency matters and duplicate reads are cheap: send a second request if the first is slow.
In practice
- Exponential backoff with full jitter
- sleep = random(0, min(cap, base × 2^attempt)).
- Retry budgets / token buckets
- Retries consume tokens that refill with successful requests.
- Circuit breakers
- Stop calling a failing dependency for a cooldown period, then probe.
It assumes
- The operation being retried is idempotent or deduplicated.
- Failures are often transient and uncorrelated with the retry itself.
- Someone observes retry rates; a rising retry rate is an early warning.
Explain it in your own words
Where you practise it
Further reading
Engineers describing it in systems they run.
- Designing robust and predictable APIs with idempotency
Stripe · Brandur Leach · Post, Feb 2017
The short, standard explanation of idempotency keys, and of retrying with backoff and jitter.
Related concepts
- Timeouts and unknown outcomes
A timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.
- Idempotency
Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.
- Backpressure and capacity
When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.
- Rate limiting
Capping how fast a client may use a resource, to protect capacity, enforce fairness, and stay within the limits of the systems you depend on.
- Delivery guarantees
At-most-once, at-least-once, and why 'exactly-once' is achieved by making duplicates harmless rather than by preventing them.
- Load shedding
Rejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.