Design a Notification System, stage 7 of 11: break it
The email provider is failing
The requirement: outages delay notifications but never lose them. Nothing about the outage is under your control except how you respond to it.
System so far· 10 parts
Select a component to see what it is responsible for and which state it owns.
- 1Product services → Event stream: Domain events via outbox
- 2Event stream → Notification planner: Events, at least once
- 3Notification planner → Notifications DB: Insert notifications under dedupe keys
- 4Notification planner → Delivery queues: Deliveries by priority class
- 5Channel senders → Delivery queues: Claim deliveries
- 6Channel senders → Email provider: Send within quota
- 7Channel senders → Push services: Send to each device
- 8Web and mobile apps → Inbox API: Inbox and read state
- 9Inbox API → Notifications DB: Read notifications
- Asynchronous
- Request / response
What you need to know
Per-message retries with backoff assume failures are independent: this message failed, the next one might not. A provider outage breaks that assumption. Every message fails for the same reason at the same time.
A circuit breaker watches failures across all calls to a dependency. When the failure rate crosses a threshold it "opens": calls stop for a while. Then it lets a few probe calls through, and closes again once they succeed.
Work it out
The email queue grows by 400 messages a second during a 25-minute outage. About how many messages are waiting when it ends?Check
Each message retries with exponential backoff and is marked failed after 5 attempts over about 2 minutes. The outage lasts 25 minutes. What happens?Failures also differ in kind, and each kind deserves a different response:
Response Kind What to do 429, 503 transient retry later, with backoff invalid address, unsubscribed permanent stop; mark failed timeout unknown a policy choice: retry and risk a duplicate, or don't