Timeouts and unknown outcomes
A timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.
Reliability
Learn it
Without a timeout, a call to a hung dependency waits forever, holding a thread, a connection and the user. With one, you stop waiting, but the request may have reached the other side and succeeded.
A timeout is a local decision to stop waiting. When it fires, the remote operation is in one of three states: never started, still running, or completed with the response lost. You can't tell which.
Check
A 'create order' request times out. What should the system record?Handling them well:
- Set timeouts deliberately: a little above the dependency's p99.9 latency, and shorter than the caller's own deadline. Propagate deadlines so inner calls give up when the outer request has.
- Reads: retry freely. Writes: retry only if idempotent; otherwise record an explicit unknown.
- Resolve unknowns by looking the operation up by its key, waiting for a callback (Webhooks), or Reconciliation.
- Bound everything: connection timeouts, read timeouts and overall deadlines are different. A slow trickle of bytes can defeat a read timeout that resets on each byte.
Think first
An API's own deadline is 2 seconds, but its call to a dependency has a 10-second timeout. What's wasted?
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Timeout treated as failure
- A succeeded charge is recorded as failed; the user pays again.
- Timeouts longer than the caller's
- Work continues for requests nobody is waiting for, wasting capacity during incidents.
- Missing timeouts
- One slow dependency exhausts every thread or connection pool upstream.
Instead, consider
- No timeout (wait indefinitely)
- Practically never in network code; acceptable only for in-process work with its own guarantees.
- Asynchronous request with status polling
- The operation is legitimately long; the caller stops waiting by design rather than by timeout.
In practice
- Per-call timeouts in HTTP clients
- Connect, read and total deadline: set all three.
- Deadline propagation (gRPC deadlines, context cancellation)
- Downstream calls inherit the remaining time.
- Unknown/pending states in data models
- Make uncertainty representable.
It assumes
- There is a way to learn the true outcome later: a lookup, a callback, or an idempotent retry.
- Callers can represent "pending" or "unknown" to users.
Explain it in your own words
Where you practise it
- A reliable video processing pipeline
- Rate limiting a public API
- Notifications across email, push and in-app
- A payment workflow that never double-charges
What the provider's behaviour implies · Should checkout wait for the provider? · Which transitions can happen? · A buyer was charged twice · Write the idempotent checkout · Thirty-seven payments stuck in processing · 100x volume and a second provider
Further reading
Engineers describing it in systems they run.
- Designing robust and predictable APIs with idempotency
Stripe · Brandur Leach · Post, Feb 2017
The short, standard explanation of idempotency keys, and of retrying with backoff and jitter.
Related concepts
- Idempotency
Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.
- Retries, backoff and jitter
Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.
- Reconciliation
Periodically comparing your records with an authoritative source and repairing differences, the backstop for every message that was lost.
- State machines for business state
Modelling an entity's lifecycle as explicit states and allowed transitions, enforced with conditional updates so concurrent or stale actors cannot corrupt it.
- Load shedding
Rejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.