Skip to content

Interactive system design interview practice

Practise architecture decisions with browser-based simulations. Change components or capacity, test failures, and see which requests fail, where work backs up, and what it takes to keep up. These exercises build reasoning between full interview rounds; they do not simulate a live interviewer.

13 simulations · Free · No account needed

Choose a system to explore

Start with a system you know, then test a change that could expose its bottleneck or failure mode.

  1. foundational · 45 min

    Design a URL Shortener

    Short links that never collide, redirect quickly for users anywhere, count every click, and can be switched off in seconds. You will estimate the load first, then design each part from the numbers.

    Scenario

    10 billion redirects and 100 million new links a month, peaks at 5× the average.

    Models: Redirects, New links, Click events, Click counts saved.

  2. intermediate · 55 min

    Design a Video Processing Pipeline

    Instructors upload multi-gigabyte lectures that take minutes to transcode. Workers crash, deploys interrupt jobs, and the same job can run twice. Every accepted upload must end in exactly one correct, visible outcome.

    Scenario

    About 2,000 uploads a day, 5× on Sunday evenings; a transcode takes about 12.5 minutes of one worker.

    Models: Upload requests, Video files stored, Transcodes, Playback requests.

  3. intermediate · 40 min

    Design an API Rate Limiter

    A public API must hold every customer to their plan across thirty stateless servers, absorb honest bursts, stop abuse, and never let the limiter itself become the outage.

    Scenario

    About 40,000 requests a second at peak across 20–60 API instances; the average is assumed.

    Models: API requests, Limiter checks, Database queries, Search queries.

  4. intermediate · 45 min

    Design a News Feed (Twitter Timeline)

    Built from Twitter's own account of its home timeline: decide when the work of 'one post, many followers' happens, survive accounts with thirty million followers, and stop precomputing timelines for people who never come back.

    Scenario

    300,000 timeline reads a second and 12,000 posts a second at peak; about 200 followers per post is assumed.

    Models: Timeline reads, Post lookups, New posts, Timeline inserts, Follower lookups.

  5. intermediate · 45 min

    Design a Distributed Job Queue

    Built from Slack's account of the outage that made it rebuild its job queue: make enqueues safe when workers fall behind, keep one slow job type from starving the rest, and drain a backlog without causing the next outage.

    Scenario

    About 1.4 billion jobs a day, peaking near 33,000 a second; 50 ms per job is assumed.

    Models: Job enqueues, Jobs run.

  6. intermediate · 45 min

    Design a Product Analytics System

    Take in billions of product events a day and answer questions nobody planned for in seconds: choose where events live, keep ingestion alive through spikes and outages, make filters on people fast, and decide what happens when two anonymous visitors turn out to be one person.

    Scenario

    About 2 billion events a day with peaks around 4× the average.

    Models: Events captured, Person updates, Dashboard queries.

  7. intermediate · 50 min

    Design a Notification System

    Product events become emails, push notifications and inbox items for millions of users. Respect every preference immediately, never notify twice, survive provider outages, and get the security alert out while a five-million-email announcement is in flight.

    Scenario

    2 million events a day become about 8 million notifications, with work-hour peaks near 10×; the email provider takes 500 a second.

    Models: Product events, Emails, Push notifications, Inbox reads.

  8. advanced · 60 min

    Design a Payment System

    Checkout calls a payment provider that can be slow, can time out after it succeeded, and sends webhooks more than once and out of order. Keep every order's payment state correct through retries, crashes and uncertainty.

    Scenario

    About 5,000 orders a day, up to 50 a minute in launches; the provider allows 100 requests a second and answers in 1.5 s.

    Models: Checkout requests, Provider charges, Webhooks, Fulfilment emails.

  9. advanced · 50 min

    Design a Distributed Cache (Memcache)

    Built from Facebook's paper on scaling memcache: a look-aside cache serving billions of reads a second, Reads are fast; the hard parts are stale values, stampedes on popular keys, dead servers and invalidations that have to cross regions.

    Scenario

    Billions of reads a second fleet-wide; scaled down here to one cluster, with a 1% miss rate assumed.

    Models: Key reads, Misses to MySQL, Writes, Invalidations.

  10. advanced · 45 min

    Design Discord's Message Storage

    Built from what Discord's engineers published about storing billions, then trillions, of messages: choose a partition key that keeps every read small, survive a channel full of deletions and a channel everyone opens at once, then move the whole thing to a new database while it is running.

    Scenario

    About 120 million messages a day with reads roughly equal to writes; the peak and ten recipients per message are assumed.

    Models: Messages sent, Message reads, Live deliveries.

  11. advanced · 60 min

    Design a Collaborative Editor (Google Docs)

    Many people edit the same document at once over unreliable connections. Every client must converge on the same text, no acknowledged keystroke may be lost, and daily deploys must not kick anyone out.

    Scenario

    20,000 connections at peak with a tenth typing 5–10 operations a second; the average and four viewers per edit are assumed.

    Models: Edits, Edits sent to viewers, Snapshots.

  12. advanced · 45 min

    Design a Reactive Database (Live Queries)

    Replace polling with live queries: work out which writes change which results, keep every screen consistent with itself, stop two people breaking a rule at the same moment, and survive one query that a whole company is watching.

    Scenario

    About 40,000 people online at peak with five live queries each, and 2,000 mutations a second.

    Models: Mutations, Query re-runs, Cached results, Invalidations.

  13. advanced · 50 min

    Shard a Live Database Without Downtime

    Built from how Notion sharded Postgres while millions of people were using it: choose a shard key, choose a shard count you can live with for years, move every row while writes continue, prove the copy is right, and grow again later without starting over.

    Scenario

    Traffic is not in the brief; these rates are assumed to show one primary running out of room.

    Models: App reads, App writes, Unmigrated writes, Change log, Backfill copies.

How the simulations work

Each playground starts with a reference architecture and an explicit workload. Select a component to change its role or capacity, put a cache or another component in front of it, or simulate it going down. Then raise traffic and compare what fails, how much work backs up, and how the changed design compares with the reference.

The workloads are interview scenarios, and some inputs are marked as assumptions. Use the results to reason about tradeoffs, not as production capacity estimates. The simulations run in your browser, and you can explore them without an account.

For a full interview exercise, start from a system's design question, make decisions before reading the answer, and use the playground to test the architecture afterward. Read the 45-minute interview guide as well. To rehearse a timed round on your own, use the system design mock interview practice.