Skip to content

Retries, backoff and jitter

Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.

Reliability

Learn it

0 of 2 checks done
  1. Many failures are transient: a dropped connection, a 503 during a deploy, a lock timeout. Retrying fixes them. Retrying badly creates new failures: synchronized waves, extra load on a struggling dependency, duplicate side effects.

    A sound retry policy answers four questions.

    1. Is it retryable? Transient errors (timeouts, 503, resets) yes. Permanent ones (400, validation, "card declined") no. Unknown outcomes only if the operation is idempotent.
    2. How long to wait? Exponential backoff (100 ms, 200 ms, 400 ms… up to a cap) gives the dependency room to recover.
    3. With jitter. If a thousand clients all wait exactly 400 ms, they all return together. Randomise the delay, e.g. "full jitter": a random value between 0 and the backoff.
    4. How many in total? Cap attempts per request, and ideally set a retry budget (retries at most ~10% of requests).
  2. Work it out

    Three layers each make up to 3 attempts at the layer below. One user action can cause up to how many calls to the bottom service?
    calls

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Retry storm
Synchronized retries without jitter hit a recovering service in waves and knock it over again.
Amplification across layers
Nested retries multiply load on the deepest dependency.
Retrying permanent errors
Wastes capacity and delays the user's real feedback.
Retrying non-idempotent calls
Duplicates charges, emails or records.

Instead, consider

Fail fast and surface the error
A human or upstream caller is better placed to decide whether to retry.
Queue for later processing
The work can wait; a durable queue retries without holding a request open.
Hedged requests
Tail latency matters and duplicate reads are cheap: send a second request if the first is slow.

In practice

Exponential backoff with full jitter
sleep = random(0, min(cap, base × 2^attempt)).
Retry budgets / token buckets
Retries consume tokens that refill with successful requests.
Circuit breakers
Stop calling a failing dependency for a cooldown period, then probe.

It assumes

  • The operation being retried is idempotent or deduplicated.
  • Failures are often transient and uncorrelated with the retry itself.
  • Someone observes retry rates; a rising retry rate is an early warning.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Timeouts and unknown outcomes

    A timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.

  • Idempotency

    Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.

  • Backpressure and capacity

    When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.

  • Rate limiting

    Capping how fast a client may use a resource, to protect capacity, enforce fairness, and stay within the limits of the systems you depend on.

  • Delivery guarantees

    At-most-once, at-least-once, and why 'exactly-once' is achieved by making duplicates harmless rather than by preventing them.

  • Load shedding

    Rejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.