Skip to content

A job queue that keeps working when workers fall behind

Design a Distributed Job Queue, from a blank page

This is how the interview actually runs: one prompt, and you decide what to cover and in what order. Write each section, then compare it with a reference design and see what you left out.

A 45-minute round. You drive; nothing prompts you.

The prompt

A team messaging app does a lot of work outside web requests: push notifications, link previews, search indexing, billing events, exports. Web servers enqueue a job and return; workers pick jobs up and run them.

Today, jobs are pushed onto lists in a set of Redis clusters, and a fleet of workers pops them off. At peak the system handles about 33,000 jobs a second, 1.4 billion a day. It has worked for years, and engineers use it for everything.

Then a database slowed down. Workers that wrote to it slowed down too, jobs piled up, and Redis ran out of memory. Dequeuing a job also needed a little free memory, so when Redis filled up the queue could not drain at all, and every feature built on jobs stopped. Slack wrote about this outage, and the system they built afterwards, in "Scaling Slack's Job Queue".

What the interviewer would tell you if you asked
  • About 1.4 billion jobs a day, peaking at about 33,000 a second.
  • Thousands of job handlers and the worker fleet talk to Redis; they cannot all be rewritten at once.
  • A Kafka cluster can be provisioned.
  • Jobs are small: a few kilobytes of JSON.
  • Most jobs finish in milliseconds; a few take seconds or minutes.
  • Jobs call downstream databases and services that have their own limits.
0:00of 45 min
0 of 5 sections written
  1. 01

    about 5 min

    What does the system have to do, and how well? List the functional requirements, then the non-functional ones (latency, availability, consistency, scale), and the questions you would ask the interviewer.

  2. 02

    about 5 min

    Turn the volumes into the numbers that drive the design: requests per second at peak, storage, bandwidth, and anything else that decides whether one machine is enough.

  3. 03

    about 10 min

    Name the components and what each one is responsible for. Then trace the main request through them, and say where the durable state lives.

  4. 04

    about 15 min

    Pick the hardest decisions in this design and make them: what you chose, what you rejected, and which constraint decided it.

  5. 05

    about 10 min

    What breaks? Walk through crashes, duplicates, slow dependencies and overload, and what the design does in each. Then: what changes at ten times the load, or with a new requirement?

Write something in at least 3 sections first. Gaps are fine; the comparison shows what they cost.