Design a Distributed Job Queue, stage 8 of 9: change it
Switch over without an outage
Slack's rollout used a shadow mode in which the relay read jobs from Kafka and discarded them instead of pushing them to Redis.
System so far· 7 parts
Select a component to see what it is responsible for and which state it owns.
- 1Web servers → Enqueue gateway: Enqueue job
- 2Enqueue gateway → Kafka: Append to topic
- 3Relay → Kafka: Read topics
- 4Relay → Redis queues: Push at a controlled rate
- 5Workers → Redis queues: Lease jobs
- 6Workers → Databases and services: Do the work
What you need to know
0 of 2 checks done
Changing a critical path safely has a standard shape:
- Build the new path and prove it is alive, with synthetic traffic.
- Run it on real traffic in parallel with the old path, with its output thrown away (shadow mode).
- Compare the two until they agree.
- Move a small, low-risk part over, with a way back.
- Move the rest in batches.
Check
Why send heartbeat jobs through every Kafka partition?Think first
In shadow mode, every job is enqueued both the old way and through the new gateway, and the relay discards what it reads. What does that let you check that synthetic tests can't?