Skip to content

Design a Distributed Job Queue, stage 8 of 9: change it

Switch over without an outage

Slack's rollout used a shadow mode in which the relay read jobs from Kafka and discarded them instead of pushing them to Redis.

System so far· 7 parts
123456SERVICEWeb serversSERVICEEnqueue gatewayLOG / STREAMKafkaWORKERRelayQUEUERedis queuesWORKERWorkersDATABASEDatabasesand services

Select a component to see what it is responsible for and which state it owns.

  1. 1Web servers → Enqueue gateway: Enqueue job
  2. 2Enqueue gateway → Kafka: Append to topic
  3. 3Relay → Kafka: Read topics
  4. 4Relay → Redis queues: Push at a controlled rate
  5. 5Workers → Redis queues: Lease jobs
  6. 6Workers → Databases and services: Do the work

What you need to know

0 of 2 checks done
  1. Changing a critical path safely has a standard shape:

    1. Build the new path and prove it is alive, with synthetic traffic.
    2. Run it on real traffic in parallel with the old path, with its output thrown away (shadow mode).
    3. Compare the two until they agree.
    4. Move a small, low-risk part over, with a way back.
    5. Move the rest in batches.
  2. Check

    Why send heartbeat jobs through every Kafka partition?