Idempotency
Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.
Reliability
Learn it
Networks drop responses, clients retry, queues redeliver, workers restart. Every retry risks doing something twice: charging, emailing, inserting. You can't reliably prevent the duplicate attempt, so you make it harmless.
An operation is idempotent if doing it twice has the same effect as once.
SET status = 'shipped'is;balance = balance - 10isn't.Check
Which of these is naturally idempotent?To make a non-idempotent operation safe, give the intent an identity:
- The caller attaches an idempotency key: a unique ID for this intent, reused on every retry.
- The server atomically claims the key:
INSERT … ON CONFLICT DO NOTHING. Whoever inserts first does the work; others find the existing record. - The server stores the outcome under the key and returns it to repeats.
Where implementations fail:
- Scope: the key identifies the intent, not the request bytes. A retried click reuses it; a new attempt (another card after a decline) gets a new one.
- In-flight duplicates: a retry can arrive while the first call runs. Return its state (202/409) rather than starting again.
- Payload mismatch: the same key with different parameters is a bug; reject it.
- Retention: keys must outlive the window in which retries arrive.
- Atomicity with the effect: record the key and perform the effect together, and pass the key downstream so external systems deduplicate too.
Think first
A client times out and retries the same charge, but generates a fresh idempotency key for the retry. What happens?
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Check-then-insert
- Two concurrent requests both see no key and both proceed. The claim must be atomic.
- Key per request, not per intent
- Generating a new key on each retry defeats the purpose.
- Key expires too soon
- A late retry, after the record has been purged, executes again.
- Effect not covered
- The key is recorded, but the side effect (an email, a downstream call) can still repeat.
Instead, consider
- Natural keys and unique constraints
- The intent already has a unique identity, such as one enrollment per (user, course).
- Conditional state transitions
- The operation is a state change that can be guarded by the current state.
- Accept duplicates and reconcile
- Duplicates are cheap and rare, and periodic cleanup is simpler than prevention.
In practice
- Idempotency-Key HTTP header
- Stripe-style: the server stores the response by key for a retention period.
- Unique constraint + ON CONFLICT
- The database performs the atomic claim.
- Message ID dedupe table
- For consumers of at-least-once queues and webhooks.
- Conditional writes (compare-and-set, ETags)
- Effects that only apply from an expected prior state.
It assumes
- Callers generate the key once per intent and reuse it across retries.
- The store that records keys supports an atomic insert-if-absent.
- Downstream systems either accept idempotency keys or their operations are naturally idempotent.
Explain it in your own words
Where you practise it
- A reliable video processing pipeline
How does the system learn the upload finished? · Two workers, one job · Defend the design
- A job queue that keeps working when workers fall behind
- Product analytics over billions of events
- Notifications across email, push and in-app
One notification per event, per channel · Trace a mention to a phone · Digests and quiet hours · Defend the duplicate policy
- A payment workflow that never double-charges
What the provider's behaviour implies · A buyer was charged twice · Write the idempotent checkout · Who owns the idempotency key? · The commit nobody acted on · Adding refunds · 100x volume and a second provider · Defend the guarantee
- A look-aside cache at Facebook's scale
- A real-time collaborative editor
Trace one keystroke · Acknowledged, then lost · Write the reconnect protocol
- Live queries: screens that update in real time
Further reading
Engineers describing it in systems they run.
- Implementing Stripe-like Idempotency Keys in Postgres
Stripe · Brandur Leach · Post, Oct 2017
The long version, with code: how to make a multi-step request safe to retry when some of its steps call other services.
- Designing robust and predictable APIs with idempotency
Stripe · Brandur Leach · Post, Feb 2017
The short, standard explanation of idempotency keys, and of retrying with backoff and jitter.
Related concepts
- Delivery guarantees
At-most-once, at-least-once, and why 'exactly-once' is achieved by making duplicates harmless rather than by preventing them.
- Retries, backoff and jitter
Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.
- State machines for business state
Modelling an entity's lifecycle as explicit states and allowed transitions, enforced with conditional updates so concurrent or stale actors cannot corrupt it.
- Concurrency control
Making read-decide-write sequences safe when other actors may change the same data in between: locks, conditional writes and constraints.
- Timeouts and unknown outcomes
A timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.
- Asynchronous processing
Accepting a request, recording the work durably, and doing it later in another process, so the work can outlive the request.
- Webhooks
HTTP callbacks from another system: delivered at least once, possibly out of order, possibly never. Handle them as hints, not truth.
- Transactions
Grouping several reads and writes so they take effect all together or not at all, isolated from concurrent work to a defined degree.
- Transactional outbox
Recording outgoing messages in the same database transaction as the state change, then delivering them separately, to avoid the dual-write problem.
- Conflict resolution and convergence
When replicas accept concurrent changes, a deterministic rule must merge them so every replica ends in the same state without losing intent.
- Leases and fencing tokens
Ownership that expires unless renewed, plus a token that lets the rest of the system reject an owner that has lost its claim without knowing it.