Skip to content

Backpressure and capacity

When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.

Performance & scale

Learn it

0 of 2 checks done
  1. Every component has a maximum throughput. When arrivals exceed it, the excess accumulates somewhere: in a queue, in buffers, in open connections, in threads waiting on locks. It's invisible until latency explodes or memory runs out, and then the whole system fails, not just the excess.

  2. Two pieces of arithmetic explain most capacity problems:

    • Little's law: items in the system = arrival rate × time each spends there (L = λW).
    • Utilization: as a server nears 100% busy, queueing delay grows roughly as 1 ÷ (1 − utilization). At 50% a request waits about one service time; at 90% about nine; at 99% about ninety.
  3. Work it out

    Requests arrive at 200 a second and each takes 50 ms. How many are in flight on average?
    requests

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Unbounded queue
Latency grows without limit; by the time work is processed, nobody wants it.
Slow consumer, fast producer
Per-connection send buffers grow until the server runs out of memory.
Retry amplification
Rejected work is retried immediately, increasing the very load that caused rejection.
Running hot
Provisioning for 95% utilization means small bursts cause large latency spikes.

Instead, consider

Add capacity (autoscaling)
Load grows predictably or slowly enough for new capacity to arrive in time.
Reduce work per request
Caching, batching or coarser updates cut the service time itself.

In practice

Bounded channels and queues
Block or reject when full.
HTTP 429/503 with Retry-After
Explicit load shedding with guidance for clients.
TCP and HTTP/2 flow control
Receivers advertise how much they can accept.
Per-connection buffer limits
Disconnect or degrade clients that fall too far behind.

It assumes

  • You can measure arrival rate, service time and queue age.
  • Upstream callers can handle rejection or slow-down signals sensibly.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Scaling Slack's Job Queue

    Slack · Saroj Yadav and others · Post, Dec 2017

    An outage, the design flaw behind it, and the redesign. The section on rolling the new path out without an outage is worth reading twice.

  • Message queues

    A durable buffer between producers and consumers that hands each message to one consumer at a time and redelivers it unless acknowledged.

  • Rate limiting

    Capping how fast a client may use a resource, to protect capacity, enforce fairness, and stay within the limits of the systems you depend on.

  • Retries, backoff and jitter

    Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.

  • Asynchronous processing

    Accepting a request, recording the work durably, and doing it later in another process, so the work can outlive the request.

  • Persistent connections

    Long-lived connections such as WebSockets turn a stateless request tier into one that holds per-client state, with consequences for routing, deploys and failure detection.

  • Partitioning

    Splitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.

  • Caching

    Keeping a copy of data closer to where it is used, trading freshness and complexity for speed and reduced load on the source.

  • Request coalescing

    When many callers ask for the same thing at the same moment, do the expensive work once and give every caller the result. The fix for thundering herds and hot keys.

  • Load shedding

    Rejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.