Webhooks
HTTP callbacks from another system: delivered at least once, possibly out of order, possibly never. Handle them as hints, not truth.
Communication
Learn it
Another system (a payment provider, a transcoder) does work for you and finishes later. A webhook lets it tell you: you register a URL, and when something happens it sends an HTTP request describing the event. Your endpoint replies
2xxto acknowledge; otherwise the provider retries with backoff, usually for hours or days, then gives up.From that contract follow the rules:
- Verify the sender. Anyone can POST to your URL. Check the signature (an HMAC of the body with a shared secret) and a timestamp to reject replays.
- Deduplicate. The same event can arrive again, even after you replied 200, if your response was lost. Record the event ID with its effect; see Delivery guarantees.
- Don't trust order. Event 2 can arrive before event 1. Apply only forward transitions (State machines for business state), or fetch the object's current state from the provider.
Check
Your endpoint processes an event and returns 200, but the response is lost on the way back. What does the provider do?And two more:
- Acknowledge fast. Persist the event and return; do slow side effects asynchronously. Slow handlers cause timeouts, timeouts cause retries, retries cause duplicates.
- Don't rely on delivery. Endpoints break, secrets rotate and providers exhaust their retries. A Reconciliation job that asks the provider about anything unresolved turns "probably" into "eventually certainly".
Think first
A secret rotation breaks signature verification for four days. The provider retries for three. What's lost, and what recovers it?
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Unverified endpoint
- An attacker POSTs a fake 'payment succeeded' event.
- Duplicate side effects
- A redelivered event provisions or emails twice.
- Regression by reordering
- A late 'processing' event overwrites a 'succeeded' state.
- Silent gap
- Webhooks stop arriving due to a bug or expired secret; nothing notices for days.
Instead, consider
- Polling the provider's API
- Volume is low, or the provider's webhooks are unreliable or unavailable.
- Provider-hosted event streams or queues
- The provider offers a pull-based event feed with cursors, which makes replay simpler.
In practice
- Inbox table
- Store event ID and payload with a unique constraint, then process asynchronously.
- Thin events + fetch
- Use the event only as a trigger, then read the authoritative state from the API.
- Reconciliation sweep
- Periodically query the provider for unresolved objects.
It assumes
- The provider signs payloads and includes a stable event ID.
- The provider offers an API to fetch the current state of an object, which reconciliation needs.
Explain it in your own words
Where you practise it
Further reading
Engineers describing it in systems they run.
- Stripe's payments APIs: the first ten years
Stripe · Michelle Bu · Post, Dec 2020
How payment methods that confirm asynchronously broke the original API, and why the replacement models a payment as one explicit state machine.
Related concepts
- Delivery guarantees
At-most-once, at-least-once, and why 'exactly-once' is achieved by making duplicates harmless rather than by preventing them.
- Reconciliation
Periodically comparing your records with an authoritative source and repairing differences, the backstop for every message that was lost.
- Idempotency
Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.
- State machines for business state
Modelling an entity's lifecycle as explicit states and allowed transitions, enforced with conditional updates so concurrent or stale actors cannot corrupt it.
- Server push: polling, long polling, SSE, WebSockets
The ways a server can tell a client that something changed, and how update rate, latency needs and who-knows-what decide between them.