Skip to content

Timeouts and unknown outcomes

A timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.

Reliability

Learn it

0 of 2 checks done
  1. Without a timeout, a call to a hung dependency waits forever, holding a thread, a connection and the user. With one, you stop waiting, but the request may have reached the other side and succeeded.

    A timeout is a local decision to stop waiting. When it fires, the remote operation is in one of three states: never started, still running, or completed with the response lost. You can't tell which.

  2. Check

    A 'create order' request times out. What should the system record?

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Timeout treated as failure
A succeeded charge is recorded as failed; the user pays again.
Timeouts longer than the caller's
Work continues for requests nobody is waiting for, wasting capacity during incidents.
Missing timeouts
One slow dependency exhausts every thread or connection pool upstream.

Instead, consider

No timeout (wait indefinitely)
Practically never in network code; acceptable only for in-process work with its own guarantees.
Asynchronous request with status polling
The operation is legitimately long; the caller stops waiting by design rather than by timeout.

In practice

Per-call timeouts in HTTP clients
Connect, read and total deadline: set all three.
Deadline propagation (gRPC deadlines, context cancellation)
Downstream calls inherit the remaining time.
Unknown/pending states in data models
Make uncertainty representable.

It assumes

  • There is a way to learn the true outcome later: a lookup, a callback, or an idempotent retry.
  • Callers can represent "pending" or "unknown" to users.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Idempotency

    Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.

  • Retries, backoff and jitter

    Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.

  • Reconciliation

    Periodically comparing your records with an authoritative source and repairing differences, the backstop for every message that was lost.

  • State machines for business state

    Modelling an entity's lifecycle as explicit states and allowed transitions, enforced with conditional updates so concurrent or stale actors cannot corrupt it.

  • Load shedding

    Rejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.