Design Discord's Message Storage, stage 4 of 9: break it
The channel that froze the cluster
Here is what the message service and the store logged when one member opened the channel. Find the lines that explain the stall.
System so far· 5 parts
Select a component to see what it is responsible for and which state it owns.
- 1Members → API servers: Send, load history, jump
- 2API servers → Gateway: New message event
- 3Gateway → Members: Push to online members
- 4API servers → Message data service: Query by channel (hash-routed)
- 5Message data service → Message cluster: Read and write one partition
- Request / response
- Asynchronous
- Server push
What you need to know
In an LSM store, a delete is a write: a tombstone marking the row as deleted. Older copies of the row may sit in files that haven't been compacted yet, so the tombstone must be kept, and read past, until compaction removes both.
Tombstones are kept for a grace period (
gc_grace_seconds) so replicas that missed the delete can learn of it through repair.Think first
A channel had 2 million messages; a bot deleted all but one. A read asks for the latest 50 messages. How much work is it?Check
A writer inserts every column, writing explicit NULL for the 12 columns a message doesn't use. What does that do in this store?Three ways to bound the damage: track which buckets are empty so reads skip them; shorten the tombstone grace period (safe if repair runs more often than the period); and stop writing needless nulls.