Stage 1 of 9 · Model
What the numbers say
Twitter reported about 30 billion timeline deliveries a day from about 400 million posts. Each timeline entry needs about 20 bytes: a post ID, the author's ID and a few flag bits.
What you need to know first
Fan-out is the multiplication of one event into many deliveries: one post, many followers' timelines. The work can happen at write time (when the post is made) or at read time (when a timeline is loaded). See Fan-out on write and fan-out on read.
Which is cheaper depends on how often each side happens. Work belongs on the rarer side.
300,000 timeline reads a second, and about 6,000 posts a second at peak. Roughly how many reads per post?
About 50 reads per post.
300,000 ÷ 6,000 = 50. Every unit of work moved from read time to write time is paid 50 times less often.
30 billion timeline deliveries a day from 400 million posts. On average, how many timelines does each post reach?
About 75 timelines.
30,000,000,000 ÷ 400,000,000 = 75. That's an average over a very uneven distribution: most accounts have hundreds of followers, a few have tens of millions.
Each user's timeline keeps 800 entries of about 20 bytes. About how many terabytes for 150 million users (one copy)?
About 2.4 TB.
800 × 20 = 16 KB per user. 16 KB × 150,000,000 = 2.4 × 10¹² bytes ≈ 2.4 TB, about 7 TB with three replicas. A big memory fleet, but an affordable one.
What the stage asks
Which statements follow?
- Holds
Timeline reads outnumber posts by roughly 50 to 1.
300,000 reads a second against about 6,000 posts a second (including peaks) is about 50:1. Work moved from reads to writes is multiplied by far fewer events.
- Holds
On average, each post is delivered to about 75 timelines.
30 billion deliveries / 400 million posts ≈ 75. That is the average; the distribution has an extremely long tail.
- Fails
Assembling timelines at read time would cost about the same as fanning out at write time.
At read time, each of 300,000 requests a second would fetch recent posts from the hundreds of accounts its reader follows and merge them: tens of millions of lookups a second plus a merge per request. At write time, ~350,000 cheap list inserts a second on average do the same job once.
- Fails
Keeping 800 post IDs for each of 150 million users needs petabytes of memory.
800 × ~20 bytes ≈ 16 KB a user; 150 million users ≈ 2.4 TB, about 7 TB with three replicas. A large in-memory fleet, but nowhere near petabytes.
The reasoning
- Reads outnumber posts about 50:1, so work moved to write time is paid far less often.
- The average post reaches about 75 timelines, but the distribution has a very long tail.
- Timelines of IDs fit in a few terabytes of memory.
Two numbers decide the design: reads are about 50 times more frequent than writes, and the average post goes to about 75 timelines. Moving the multiplication to write time is the cheaper side for the average post, and keeping the result in memory is affordable.
The word average is doing a lot of work there. The tail of that distribution is the subject of stage 4.