Asynchronous processing
Accepting a request, recording the work durably, and doing it later in another process, so the work can outlive the request.
Communication
Learn it
Some work doesn't fit inside a request: a 15-minute video transcode, a large report, a call to a slow partner API. Some work just doesn't need to finish before you answer, like sending a receipt email.
If that work runs inside the request, a deploy, crash or timeout loses it, and the client can't tell whether it happened.
Doing it asynchronously: the request records the work durably (a row in a jobs table or a message in a queue) and responds
202 Acceptedwith an id. A separate worker picks up the work and does it later. The client checks the status with the id, or is notified when it's done.Check
When is it safe to respond 202 Accepted?Think first
A worker is halfway through a job when its machine dies. What has to exist so that the job isn't stuck as "processing" forever?Check
Because of that recovery, what must be true of the job?Check
Does moving work to a queue make the work finish faster?One thing to monitor: if jobs arrive faster than workers finish them, nothing errors. The queue just grows, and work gets later and later. Alert on the age of the oldest waiting job, since that's what users experience.
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Accepted but not recorded
- The API responds 202 and enqueues the work in memory or after the response. A crash in between loses work the client was told was accepted.
- Stranded work
- A worker dies mid-job and nothing reassigns it; the status says processing forever.
- Silent backlog
- Work arrives faster than workers finish it. Nothing errors; latency just grows until users notice. Queue age, not queue length, is the metric to alert on.
- Duplicate side effects
- Retried work charges, emails or writes twice because it was not designed to run more than once.
Instead, consider
- Do it inside the request
- The work reliably finishes well within the request deadline and losing it on a crash is acceptable, because the client will retry.
- Streaming the response
- The work is long but the client must watch it happen and can stay connected, as in a build log or an LLM response.
- Scheduled batch processing
- Results are needed periodically rather than per request, and throughput matters more than latency.
In practice
- Background thread in the same process
- Not durable: fine only for work you are happy to lose.
- Jobs table in the primary database
- Durable and transactional with your data; good up to moderate throughput.
- Managed queue (SQS, Cloud Tasks, etc.)
- Durable delivery with visibility timeouts and dead-letter handling built in.
- Workflow engines (Temporal, Step Functions)
- For multi-step processes that need durable state between steps.
It assumes
- The caller can tolerate not knowing the result when the request returns.
- The accepted work is recorded durably before the response is sent; otherwise "accepted" is a lie.
- There is a way for the caller to learn the eventual outcome.
Explain it in your own words
Where you practise it
Further reading
Engineers describing it in systems they run.
- How to transcode video 100x faster; or, a Gordian knot cut
Mux · Jon Dahl · Post, Apr 2023
Why transcoding everything before publishing is slow, and the alternative: split the upload into segments and transcode each one the first time someone watches it.
- Dub's link redirect middleware
Dub · Steven Tey and Dub contributors · Code, Sep 2022
The code that serves every Dub short link. Look for the Redis lookup with a database fallback, the click recorded after the response is sent, and what happens when Redis is failing over.
- Stripe's payments APIs: the first ten years
Stripe · Michelle Bu · Post, Dec 2020
How payment methods that confirm asynchronously broke the original API, and why the replacement models a payment as one explicit state machine.
Related concepts
- Message queues
A durable buffer between producers and consumers that hands each message to one consumer at a time and redelivers it unless acknowledged.
- Idempotency
Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.
- Leases and fencing tokens
Ownership that expires unless renewed, plus a token that lets the rest of the system reject an owner that has lost its claim without knowing it.
- State machines for business state
Modelling an entity's lifecycle as explicit states and allowed transitions, enforced with conditional updates so concurrent or stale actors cannot corrupt it.
- Backpressure and capacity
When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.
- Object storage
A flat namespace of immutable blobs addressed by key, built for durability and size rather than queries or in-place updates.
- Fan-out on write and fan-out on read
When one write must reach many readers, do the work when it is written (precompute every reader's view) or when it is read (assemble it on demand). Most real feeds do both.