Stage 1 of 12 · Model
Size the problem
Real-time systems fail in ways that depend on numbers: message rates, fan-out factors, connection counts. Work some out before you choose anything. Useful figures: 20,000 connections at peak, about a tenth of users typing at any moment, 5-10 operations per second per typist, and one document with 30 editors and 200 viewers.
What you need to know first
"Last write wins" on the whole document means: whoever saves second replaces whatever the first person wrote. With two people typing in the same second, someone's sentence disappears.
Collaborative editing has to merge operations (insert "x" at this point, delete these characters), not replace whole documents.
20,000 people connected, a tenth of them typing at about 7 operations a second. About how many operations a second arrive?
About 14,000 ops per second.
20,000 × 0.1 × 7 = 14,000 ops a second. As 14,000 separate committed transactions that's heavy; batched per document it's modest.
The all-hands doc: 30 editors at about 7 ops a second each, every op delivered to about 230 participants. About how many outbound messages a second?
About 46,000 messages per second.
30 × 7 ≈ 210 ops a second; × 230 recipients ≈ 48,000 messages a second (about 46,000, excluding each sender) for one document. Delivery, not ingestion, is the hot path.
Do 20,000 idle WebSocket connections need a large server fleet?
No: an idle connection costs tens of kilobytes on an event-driven server; what costs is message rate and buffering.
20,000 × ~30 KB is well under a gigabyte. A few servers hold the connections; traffic decides the rest.
What the stage asks
Which statements hold?
- Fails
Sending the whole document on each change works if the change is debounced to once a second.
Bandwidth is the smaller problem. Whole-document last-write-wins means any two people typing within the same second overwrite each other, which is exactly the bug in testing. The unit of change has to be an operation, merged rather than replaced; see Conflict resolution and convergence.
- Holds
Polling once a second cannot meet the ~200 ms latency goal.
Average delay would be half a second plus request time, and polling faster multiplies requests for every idle client.
- Holds
The all-hands document alone can require tens of thousands of outbound messages per second.
30 editors × ~7 ops/s ≈ 200 ops/s, each delivered to ~230 participants, is about 46,000 messages a second for one document, unless you batch. Fan-out, not ingestion, is the hot path; see Backpressure and capacity.
- Depends
20,000 concurrent WebSocket connections require a large server fleet.
Idle connections are cheap on an event-driven server: tens of kilobytes each, so 20,000 fit in well under a gigabyte. What costs is message rate and per-connection buffering. A handful of servers holds the connections; the fleet size is set by ops and fan-out. See Persistent connections.
- Depends
Writing every operation durably to Postgres is infeasible at this scale.
About 2,000 typists × ~7 ops/s ≈ 14,000 ops a second. As 14,000 separate committed transactions, that is a heavy load for one primary. Batched per document, committing every 10-20 ms with the ack waiting for the batch, it becomes a few hundred transactions a second carrying small rows, which is very manageable.
The reasoning
- Concurrent edits are normal: merge operations instead of replacing documents.
- About 14,000 ops a second inbound means writes must be batched.
- One hot document can need tens of thousands of outbound messages a second; connections themselves are cheap.
Three numbers frame the design:
- ~14,000 ops/s inbound, so writes must be batched (without weakening acknowledgements).
- ~46,000 msgs/s for one hot document, so fan-out needs batching and its own capacity.
- 20,000 connections, which is less than it sounds. Connections are not the bottleneck; traffic is.
And one non-number: concurrent edits are normal, not an edge case. The design has to merge rather than replace.