Persistent connections
Long-lived connections such as WebSockets turn a stateless request tier into one that holds per-client state, with consequences for routing, deploys and failure detection.
Communication
Learn it
A stateless HTTP server forgets each client after responding, so any server can handle any request and deploys are trivial. A long-lived connection (a WebSocket) changes that: a specific server now holds a specific client for minutes or days, with buffers, subscriptions and session state.
Holding connections brings responsibilities:
- Liveness. TCP doesn't tell you promptly that the other side vanished. A laptop lid closing sends nothing. Applications send heartbeats every 15–30 s and treat a few missed ones as dead.
- Routing. Messages for a client must reach the server holding its connection: route related clients together (Partitioning by room or document), or connect servers with Publish/subscribe.
- Resumption. Connections drop constantly on mobile networks. The client reconnects, maybe elsewhere, and needs a position (a sequence number or last event ID) to resume from.
Check
A phone's battery dies mid-session. When does the server notice?Two more:
- Deploys. Draining doesn't end long-lived connections. Servers must tell clients to reconnect, with jittered delays so the new fleet isn't hit by every client at once.
- Flow control. A slow client's outbound buffer grows on the server. Without a bound, one slow client can exhaust memory; see Backpressure and capacity.
Connection count alone is usually cheap: an idle WebSocket costs kilobytes. The cost is message rate, buffers, and the coordination above.
Work it out
50,000 idle WebSockets at about 30 KB each. About how many megabytes of memory?
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Ghost connections
- Without heartbeats, dead clients look connected and receive messages into a void.
- Thundering herd on deploy
- Every client reconnects at the same instant and overloads the remaining servers.
- Lost messages across reconnects
- No resume position means anything sent during the gap is gone.
- Unbounded send buffers
- A slow reader causes the server to queue its messages until memory runs out.
Instead, consider
- Polling or SSE
- Updates are infrequent or one-directional and statelessness is worth more than latency.
- Managed realtime service
- You want the protocol benefits without operating connection-holding servers.
In practice
- WebSocket server with heartbeat
- ws, uWebSockets, Go's gorilla/nhooyr libraries; you build resumption.
- Socket.IO and similar
- Adds reconnection, rooms and fallbacks on top of WebSockets.
- Managed realtime platforms
- Ably, Pusher, Firebase and others: hosted connections and fan-out.
It assumes
- Load balancers and proxies support upgrade and long idle times, or heartbeats keep connections from being cut.
- Clients implement reconnection with backoff and resumption.
Explain it in your own words
Where you practise it
Further reading
Engineers describing it in systems they run.
- Uber's Real-Time Push Platform
Uber · Madan Thangavelu and others · Post, Dec 2020
Why polling was replaced, and the delivery problems push brought with it: resuming after a dropped connection, and knowing what actually arrived.
- How Figma's multiplayer technology works
Figma · Evan Wallace · Post, Oct 2019
Why a central server per document let Figma avoid most of the machinery of operational transforms, and how ordering works when anyone can insert anywhere.
Related concepts
- Server push: polling, long polling, SSE, WebSockets
The ways a server can tell a client that something changed, and how update rate, latency needs and who-knows-what decide between them.
- Publish/subscribe
Decoupling senders from receivers by topic: a publisher sends once and every current subscriber receives a copy.
- Backpressure and capacity
When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.
- Partitioning
Splitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.
- Soft state
State that expires unless refreshed. It is cheap to keep, safe to lose, and right for presence, sessions and anything that describes the present moment.