Interactive system design interview practice
Practise architecture decisions with browser-based simulations. Change components or capacity, test failures, and see which requests fail, where work backs up, and what it takes to keep up. These exercises build reasoning between full interview rounds; they do not simulate a live interviewer.
13 simulations · Free · No account needed
Choose a system to explore
Start with a system you know, then test a change that could expose its bottleneck or failure mode.
foundational · 45 min
Design a URL Shortener
Short links that never collide, redirect quickly for users anywhere, count every click, and can be switched off in seconds. You will estimate the load first, then design each part from the numbers.
Scenario
10 billion redirects and 100 million new links a month, peaks at 5× the average.
Models: Redirects, New links, Click events, Click counts saved.
intermediate · 55 min
Design a Video Processing Pipeline
Instructors upload multi-gigabyte lectures that take minutes to transcode. Workers crash, deploys interrupt jobs, and the same job can run twice. Every accepted upload must end in exactly one correct, visible outcome.
Scenario
About 2,000 uploads a day, 5× on Sunday evenings; a transcode takes about 12.5 minutes of one worker.
Models: Upload requests, Video files stored, Transcodes, Playback requests.
intermediate · 40 min
Design an API Rate Limiter
A public API must hold every customer to their plan across thirty stateless servers, absorb honest bursts, stop abuse, and never let the limiter itself become the outage.
Scenario
About 40,000 requests a second at peak across 20–60 API instances; the average is assumed.
Models: API requests, Limiter checks, Database queries, Search queries.
intermediate · 45 min
Design a News Feed (Twitter Timeline)
Built from Twitter's own account of its home timeline: decide when the work of 'one post, many followers' happens, survive accounts with thirty million followers, and stop precomputing timelines for people who never come back.
Scenario
300,000 timeline reads a second and 12,000 posts a second at peak; about 200 followers per post is assumed.
Models: Timeline reads, Post lookups, New posts, Timeline inserts, Follower lookups.
intermediate · 45 min
Design a Distributed Job Queue
Built from Slack's account of the outage that made it rebuild its job queue: make enqueues safe when workers fall behind, keep one slow job type from starving the rest, and drain a backlog without causing the next outage.
Scenario
About 1.4 billion jobs a day, peaking near 33,000 a second; 50 ms per job is assumed.
Models: Job enqueues, Jobs run.
intermediate · 45 min
Design a Product Analytics System
Take in billions of product events a day and answer questions nobody planned for in seconds: choose where events live, keep ingestion alive through spikes and outages, make filters on people fast, and decide what happens when two anonymous visitors turn out to be one person.
Scenario
About 2 billion events a day with peaks around 4× the average.
Models: Events captured, Person updates, Dashboard queries.
intermediate · 50 min
Design a Notification System
Product events become emails, push notifications and inbox items for millions of users. Respect every preference immediately, never notify twice, survive provider outages, and get the security alert out while a five-million-email announcement is in flight.
Scenario
2 million events a day become about 8 million notifications, with work-hour peaks near 10×; the email provider takes 500 a second.
Models: Product events, Emails, Push notifications, Inbox reads.
advanced · 60 min
Design a Payment System
Checkout calls a payment provider that can be slow, can time out after it succeeded, and sends webhooks more than once and out of order. Keep every order's payment state correct through retries, crashes and uncertainty.
Scenario
About 5,000 orders a day, up to 50 a minute in launches; the provider allows 100 requests a second and answers in 1.5 s.
Models: Checkout requests, Provider charges, Webhooks, Fulfilment emails.
advanced · 50 min
Design a Distributed Cache (Memcache)
Built from Facebook's paper on scaling memcache: a look-aside cache serving billions of reads a second, Reads are fast; the hard parts are stale values, stampedes on popular keys, dead servers and invalidations that have to cross regions.
Scenario
Billions of reads a second fleet-wide; scaled down here to one cluster, with a 1% miss rate assumed.
Models: Key reads, Misses to MySQL, Writes, Invalidations.
advanced · 45 min
Design Discord's Message Storage
Built from what Discord's engineers published about storing billions, then trillions, of messages: choose a partition key that keeps every read small, survive a channel full of deletions and a channel everyone opens at once, then move the whole thing to a new database while it is running.
Scenario
About 120 million messages a day with reads roughly equal to writes; the peak and ten recipients per message are assumed.
Models: Messages sent, Message reads, Live deliveries.
advanced · 60 min
Design a Collaborative Editor (Google Docs)
Many people edit the same document at once over unreliable connections. Every client must converge on the same text, no acknowledged keystroke may be lost, and daily deploys must not kick anyone out.
Scenario
20,000 connections at peak with a tenth typing 5–10 operations a second; the average and four viewers per edit are assumed.
Models: Edits, Edits sent to viewers, Snapshots.
advanced · 45 min
Design a Reactive Database (Live Queries)
Replace polling with live queries: work out which writes change which results, keep every screen consistent with itself, stop two people breaking a rule at the same moment, and survive one query that a whole company is watching.
Scenario
About 40,000 people online at peak with five live queries each, and 2,000 mutations a second.
Models: Mutations, Query re-runs, Cached results, Invalidations.
advanced · 50 min
Shard a Live Database Without Downtime
Built from how Notion sharded Postgres while millions of people were using it: choose a shard key, choose a shard count you can live with for years, move every row while writes continue, prove the copy is right, and grow again later without starting over.
Scenario
Traffic is not in the brief; these rates are assumed to show one primary running out of room.
Models: App reads, App writes, Unmigrated writes, Change log, Backfill copies.
How the simulations work
Each playground starts with a reference architecture and an explicit workload. Select a component to change its role or capacity, put a cache or another component in front of it, or simulate it going down. Then raise traffic and compare what fails, how much work backs up, and how the changed design compares with the reference.
The workloads are interview scenarios, and some inputs are marked as assumptions. Use the results to reason about tradeoffs, not as production capacity estimates. The simulations run in your browser, and you can explore them without an account.
For a full interview exercise, start from a system's design question, make decisions before reading the answer, and use the playground to test the architecture afterward. Read the 45-minute interview guide as well. To rehearse a timed round on your own, use the system design mock interview practice.