Product analytics over billions of events
Design a Product Analytics System, from a blank page
This is how the interview actually runs: one prompt, and you decide what to cover and in what order. Write each section, then compare it with a reference design and see what you left out.
A 45-minute round. You drive; nothing prompts you.
The prompt
An analytics product for software teams. Customers add a snippet to their app, and it sends events: a page was viewed, a button clicked, a plan upgraded. Each event has a name, a timestamp, the ID of whoever did it, and a bag of properties (browser, country, plan, and whatever else the customer's developers decide to send).
Customers build charts from these events without writing SQL: daily signups from Germany, the share of visitors who sign up and then pay within a week, how many people come back a month later. Each chart is a new question; the product cannot know in advance which ones will be asked.
There are about 6,000 customer teams, sending about 2 billion events a day in total, and the biggest few teams send a large share of it. Everything currently lives in one Postgres table with the properties in a JSON column. Charts for small teams are fine. Charts for big teams time out.
What the interviewer would tell you if you asked
- About 2 billion events a day, with peaks around four times the average.
- The largest teams each send hundreds of millions of events a month.
- Properties are free-form: thousands of different keys across all customers.
- An event is about 1 KB as sent, most of it properties.
- Queries almost always filter by one team and a time range.
- Events are never edited after they arrive; people's properties change over time.
01
about 5 minWhat does the system have to do, and how well? List the functional requirements, then the non-functional ones (latency, availability, consistency, scale), and the questions you would ask the interviewer.
02
about 5 minTurn the volumes into the numbers that drive the design: requests per second at peak, storage, bandwidth, and anything else that decides whether one machine is enough.
03
about 10 minName the components and what each one is responsible for. Then trace the main request through them, and say where the durable state lives.
04
about 15 minPick the hardest decisions in this design and make them: what you chose, what you rejected, and which constraint decided it.
05
about 10 minWhat breaks? Walk through crashes, duplicates, slow dependencies and overload, and what the design does in each. Then: what changes at ten times the load, or with a new requirement?
Write something in at least 3 sections first. Gaps are fine; the comparison shows what they cost.