Skip to content

Storing trillions of chat messages

Design Discord's Message Storage, from a blank page

This is how the interview actually runs: one prompt, and you decide what to cover and in what order. Write each section, then compare it with a reference design and see what you left out.

A 45-minute round. You drive; nothing prompts you.

The prompt

A chat platform organises conversations into servers, and servers into channels. Opening a channel loads its latest 50 messages; scrolling up loads older pages; clicking a reply or a search result jumps straight to one message. Messages can be edited and deleted.

The platform stores about 120 million messages a day and the number is climbing fast. Reads and writes are roughly equal, and reads are scattered: most servers are small groups of friends, a few are public communities with hundreds of thousands of members. The current database is a single replica set whose data and indexes no longer fit in memory, and read latency has become unpredictable.

Discord described this exact situation in 2017, and what happened over the five years after it in 2023.

What the interviewer would tell you if you asked
  • About 120 million new messages a day, growing several-fold a year.
  • Roughly equal reads and writes; reads are spread randomly across millions of channels.
  • A small infrastructure team: operating the store must not need constant manual work.
  • A message is about 1 KB including metadata.
  • Message IDs are 64-bit Snowflakes: a millisecond timestamp, a worker number and a sequence, so they sort by time.
  • New messages reach online members through a separate WebSocket gateway; this investigation is about storing and reading them.
0:00of 45 min
0 of 5 sections written
  1. 01

    about 5 min

    What does the system have to do, and how well? List the functional requirements, then the non-functional ones (latency, availability, consistency, scale), and the questions you would ask the interviewer.

  2. 02

    about 5 min

    Turn the volumes into the numbers that drive the design: requests per second at peak, storage, bandwidth, and anything else that decides whether one machine is enough.

  3. 03

    about 10 min

    Name the components and what each one is responsible for. Then trace the main request through them, and say where the durable state lives.

  4. 04

    about 15 min

    Pick the hardest decisions in this design and make them: what you chose, what you rejected, and which constraint decided it.

  5. 05

    about 10 min

    What breaks? Walk through crashes, duplicates, slow dependencies and overload, and what the design does in each. Then: what changes at ten times the load, or with a new requirement?

Write something in at least 3 sections first. Gaps are fine; the comparison shows what they cost.