Skip to content

Persistent connections

Long-lived connections such as WebSockets turn a stateless request tier into one that holds per-client state, with consequences for routing, deploys and failure detection.

Communication

Learn it

0 of 2 checks done
  1. A stateless HTTP server forgets each client after responding, so any server can handle any request and deploys are trivial. A long-lived connection (a WebSocket) changes that: a specific server now holds a specific client for minutes or days, with buffers, subscriptions and session state.

  2. Holding connections brings responsibilities:

    • Liveness. TCP doesn't tell you promptly that the other side vanished. A laptop lid closing sends nothing. Applications send heartbeats every 15–30 s and treat a few missed ones as dead.
    • Routing. Messages for a client must reach the server holding its connection: route related clients together (Partitioning by room or document), or connect servers with Publish/subscribe.
    • Resumption. Connections drop constantly on mobile networks. The client reconnects, maybe elsewhere, and needs a position (a sequence number or last event ID) to resume from.
  3. Check

    A phone's battery dies mid-session. When does the server notice?

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Ghost connections
Without heartbeats, dead clients look connected and receive messages into a void.
Thundering herd on deploy
Every client reconnects at the same instant and overloads the remaining servers.
Lost messages across reconnects
No resume position means anything sent during the gap is gone.
Unbounded send buffers
A slow reader causes the server to queue its messages until memory runs out.

Instead, consider

Polling or SSE
Updates are infrequent or one-directional and statelessness is worth more than latency.
Managed realtime service
You want the protocol benefits without operating connection-holding servers.

In practice

WebSocket server with heartbeat
ws, uWebSockets, Go's gorilla/nhooyr libraries; you build resumption.
Socket.IO and similar
Adds reconnection, rooms and fallbacks on top of WebSockets.
Managed realtime platforms
Ably, Pusher, Firebase and others: hosted connections and fan-out.

It assumes

  • Load balancers and proxies support upgrade and long idle times, or heartbeats keep connections from being cut.
  • Clients implement reconnection with backoff and resumption.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Uber's Real-Time Push Platform

    Uber · Madan Thangavelu and others · Post, Dec 2020

    Why polling was replaced, and the delivery problems push brought with it: resuming after a dropped connection, and knowing what actually arrived.

  • How Figma's multiplayer technology works

    Figma · Evan Wallace · Post, Oct 2019

    Why a central server per document let Figma avoid most of the machinery of operational transforms, and how ordering works when anyone can insert anywhere.

  • Server push: polling, long polling, SSE, WebSockets

    The ways a server can tell a client that something changed, and how update rate, latency needs and who-knows-what decide between them.

  • Publish/subscribe

    Decoupling senders from receivers by topic: a publisher sends once and every current subscriber receives a copy.

  • Backpressure and capacity

    When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.

  • Partitioning

    Splitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.

  • Soft state

    State that expires unless refreshed. It is cheap to keep, safe to lose, and right for presence, sessions and anything that describes the present moment.