Skip to content

The finished design, decision by decision

How to design a Notification System

Not the one correct diagram, but a design you can defend under these constraints: the finished architecture, then every stage's question with the reasoning that answers it, the tradeoffs it accepts, and where another engineer could land differently.

Notifications across email, push and in-app

The short answer

10 parts, each with one job. The map below shows how requests and data move between them; the stages after it explain why each part is there.

Product services
Write domain events to their outbox in the same transaction as the action.
Event stream
Relays outbox events to consumers, at least once.
Notification planner
Turns an event into notifications: resolves recipients, applies preferences and quiet hours, inserts one row per user and channel under a dedupe key.
Notifications DB
Notifications, deliveries and preferences. The source of truth for what was sent and what the inbox shows.
Delivery queues
Separate queues per priority class (critical, transactional, bulk) and channel.
Channel senders
Claim a delivery, re-check preferences, take a token from the class's provider quota, send, and record the outcome.
Email provider
Delivers email; 500/s account limit.
Push services
APNs and FCM; report invalid tokens.
Inbox API
Serves the in-app inbox and read state from the notifications DB.
Web and mobile apps
Show the inbox and receive push.
123456789SERVICEProduct servicesLOG / STREAMEvent streamWORKERNotificationplannerDATABASENotifications DBQUEUEDelivery queuesWORKERChannel sendersEXTERNALEmail providerEXTERNALPush servicesSERVICEInbox APICLIENTWeb andmobile apps

Select a component to see what it is responsible for and which state it owns.

  1. 1Product services → Event stream: Domain events via outbox
  2. 2Event stream → Notification planner: Events, at least once
  3. 3Notification planner → Notifications DB: Insert notifications under dedupe keys
  4. 4Notification planner → Delivery queues: Deliveries by priority class
  5. 5Channel senders → Delivery queues: Claim deliveries
  6. 6Channel senders → Email provider: Send within quota
  7. 7Channel senders → Push services: Send to each device
  8. 8Web and mobile apps → Inbox API: Inbox and read state
  9. 9Inbox API → Notifications DB: Read notifications
  • Asynchronous
  • Request / response

Why does this design work?

Product actions emit events, durably and atomically, and never wait for delivery. A planner turns events into notifications with derived identities, so replays cannot multiply them. Each delivery is a state machine claimed with a token, re-checked against current preferences at the last moment, metered by a per-class quota, and classified on failure as transient, permanent or unknown.

Isolation comes from capacity rather than ordering: critical traffic has its own queue and reserved share of the provider limit. Duplicates are structurally impossible up to the provider boundary, and beyond it they are a deliberate, per-type trade between duplicate and silence.

Invariants, and where they are enforced

  • At most one notification exists per (user, event, channel).

    A unique dedupe key built from those three parts; redelivered events hit the constraint and create nothing. Enforced by Notifications DB, Notification planner.

  • Nothing is sent that the user's current preferences forbid.

    Preferences are applied when planning and re-checked immediately before each provider call. Enforced by Channel senders, Notification planner.

  • Bulk traffic cannot delay critical notifications.

    Separate queues and a reserved slice of the provider quota per priority class. Enforced by Delivery queues, Channel senders.

  • A planned delivery eventually succeeds, fails permanently with a reason, or is cancelled by preference.

    Delivery rows with a state machine; retries with backoff for transient errors; permanent errors terminate. Enforced by Notifications DB, Channel senders.

What does it rely on?

  • Product services write events through an outbox, with stable event ids.
  • The notifications database enforces dedupe keys and conditional transitions.
  • Provider rate limits are known and enforced by a shared limiter across senders.
  • Queue leases (visibility timeouts) are longer than the slowest send, or are extended.

What tradeoffs does it make?

ChoiceGainsCosts
Event-driven decouplingProduct actions never fail because of notifications.Notifications are eventually sent, seconds later, and the pipeline needs monitoring.
Send-time preference checksChanges take effect immediately.A preference read per send.
Reserved quota per classCritical alerts meet their deadline under any bulk load.Bulk sends finish later; reserved capacity sometimes idles.
No retry on unknown outcomes for routine mailFewer duplicates.Occasional missed emails, mitigated by the inbox.

What are the reasonable alternatives?

Notification-as-a-service platform
Better when channels, templates and preferences are standard and you would rather configure than build. You still own the event boundary and idempotency.
Per-channel microservices
Better when channels have very different scale or teams; a shared planner still owns identity and policy.
Inbox-first (pull) model
Better when most notifications are low-urgency: write to the inbox only and send email digests on a schedule, which removes most real-time sends.

When does it stop working?

  • Providers change their limits or semantics without notice, so quotas need monitoring and configuration, not constants.
  • A single event fans out to millions of recipients, such as a workspace-wide announcement, so planning itself must be batched and paged.
  • Regulatory requirements demand a proof of delivery the providers cannot give.
  • Users expect cross-device read sync in real time, which turns the inbox into a real-time sync problem.

Every stage, decided and explained

Spoilers, for the whole investigation: each stage's question and its answer, the reasoning behind it, and the tradeoffs it accepts. If you have not worked through the stages yet, you may want to do that first.

Work through the stages

Stage 1 of 11 · Model

What the numbers say

Start with arithmetic and the provider's contract. Eight million notifications a day, a 500/s email limit, a five-million-user announcement, and a 30-second promise for security alerts.

What you need to know first

When a provider caps your sending rate, that cap is a shared budget. Every email of every kind draws from it: security alerts, comment notifications and marketing alike. The first numbers to work out are how long the big jobs occupy that budget, and what is left for everything else meanwhile.

An announcement goes to 5 million users, and the provider allows 500 emails a second. About how many hours does it take?

About 2.8 hours.

5,000,000 ÷ 500 = 10,000 seconds ≈ 2.8 hours. For that whole time, the account has no spare capacity unless something reserves it.

8 million notifications a day, with work-hour peaks at 10× the average. About how many a second at peak?

About 930 per second.

8,000,000 ÷ 86,400 ≈ 93 a second on average. × 10 ≈ 930 a second at peak.

If most of them are emails, the peak alone is nearly twice the provider's 500 a second. Something has to wait, and the design decides what.

An idempotency key is an id you send with a request so the receiver can recognise a retry: "you already did 812, here's the same result". Payment APIs usually accept one. This email provider doesn't.

Without it, a request that times out is ambiguous: the email may or may not have been sent, and the provider gives you no way to ask.

A send to the provider times out after 10 seconds. What do you know?

Nothing for certain: it may have been sent, or not.

The timeout is on your side. The provider may have sent the email and its response got lost, or it may never have received the request. Retrying risks a duplicate; not retrying risks silence.

What the stage asks

Which statements hold?

  1. Holds

    Sending a marketing email to all 5 million users takes nearly three hours at the provider's limit.

    5,000,000 / 500 per second = 10,000 seconds, about 2.8 hours. For that long, the account's entire send capacity is in use.

  2. Holds

    If security alerts share a FIFO queue with the announcement, they can wait hours.

    A security alert enqueued just after the announcement waits behind millions of messages draining at 500/s. The 30-second promise requires isolation from bulk traffic, both in queueing and in quota.

  3. Fails

    Sending notifications inside the request that created the comment is fine, since each comment notifies only a few people.

    The count is small but the dependency is the problem: provider latency becomes comment latency, and provider outages become comment failures, which is last week's incident. The product action must succeed without the notification having been sent.

  4. Fails

    With careful engineering, every email can be guaranteed to be delivered exactly once.

    The provider accepts no idempotency key. If a send times out, you cannot know whether it went out; retrying risks a duplicate and not retrying risks silence. You can make duplicates rare, but not impossible. See Delivery guarantees.

  5. Fails

    Eight million notifications a day is about 93 a second, so peak capacity is not a concern.

    Averages hide peaks. Work hours run around 10x the average, roughly 900 a second, which already exceeds the email limit if most notifications are email. Size for peaks, and decide what waits.

The reasoning

  1. A provider's rate limit is a shared budget; bulk sends can occupy all of it for hours.
  2. Size for peaks, not averages, and decide in advance what waits when demand exceeds the limit.
  3. Without provider idempotency keys, a timed-out send is ambiguous, so exactly-once email is impossible.

The provider's limit makes email capacity a shared, scarce resource, and the announcement can consume all of it for hours. Most of the design is about who gets that capacity and when.

The missing idempotency key is the other defining fact: for email, "exactly once" is off the table, so the goal becomes at most once per notification, with the residual duplicate window as small as possible.

Stage 2 of 11 · Decide

Decouple from the product action

Someone comments on an issue. Three followers should be notified. The comment must save quickly and reliably whatever the email provider is doing.

What you need to know first

When a request does two things, the slower and less reliable one sets the pace for both. If saving a comment also sends emails, the comment is as slow as the email provider and fails whenever the provider fails.

The fix is to make the comment record that notifications are owed, and let something else do the sending later.

After saving the comment, the server starts an in-memory background task to send the emails, then responds. A deploy restarts the server 50 ms later. What happens?

The task dies with the process, and the emails are never sent. Nothing records that they were owed, so nothing retries them, and nobody notices.

The intent to notify has to be stored somewhere durable before the request returns.

A dual write updates two systems separately, say the database and then a queue. A crash between the two leaves them disagreeing: a comment with no notification, or a notification for a comment that was rolled back.

The transactional outbox avoids it: write the comment and an event row ("comment 991 created") in the same database transaction. Either both exist or neither does. A relay later reads new outbox rows and publishes them. See Transactional outbox.

With an outbox, the comment and its event commit together. The relay crashes before publishing the event. What happens?

The event is still in the outbox table, and the relay publishes it when it restarts.

The event is durable from the moment the comment committed. Publishing can be late, but it can't be lost.

What the stage asks

How should a comment lead to notifications?

  1. Flawed

    The comment endpoint sends the notifications before responding

    This couples the comment's latency and availability to every provider, which was last week's incident.

  2. Flawed

    After saving the comment, start a background task in the same process to send notifications

    The comment is fast now, but the background task dies with the process. A deploy right after the commit silently drops notifications, with no record that any were owed.

  3. Sound

    Save the comment and a comment-created event in one transaction; a notification planner consumes events asynchronously

    The comment succeeds whenever the database does. The event is durable from the same commit, so notifications are owed even if every process dies a millisecond later. Product services know nothing about channels, preferences or providers; the notification system owns all of that.

  4. Defensible

    Call a notification service's API synchronously, which queues the work internally

    Better than sending inline, because the service queues the work. But the comment still fails when the notification service is down, and a crash between saving the comment and calling the service loses the notification. It is a dual write, which the outbox exists to remove.

What a strong answer covers

  • The product action's success and latency must not depend on notification delivery.
  • The intent to notify must be recorded durably and atomically with the action (outbox/event).
  • Channel, preference and provider logic belong to the notification system, not each product service.Supporting

The reasoning

  1. The product action must not depend on notification delivery for its latency or success.
  2. Record the intent to notify atomically with the action (an outbox event), then deliver asynchronously.
  3. Product services emit facts; the notification system owns channels, preferences and providers.

This is the Transactional outbox used as an architectural boundary. Product services emit facts ("comment 991 was created"); the notification system decides what those facts mean for whom. Adding a channel, changing a template or switching providers then touches one system, not twenty.

The event stream delivers at least once, so the planner must tolerate seeing the same event twice. That is the next problem.

Stage 3 of 11 · Decide

One notification per event, per channel

The event stream redelivers events after consumer restarts. Two planner instances can process the same event concurrently. Each must produce the same notifications, once.

What you need to know first

Event streams deliver at least once. A consumer that processes an event and crashes before recording its progress will see that event again after restarting. Two consumer instances can also briefly process the same event during a rebalance.

So the planner must produce the same result however many times it sees an event.

The most robust way is to give each notification an identity derived from its cause: (recipient, event id, channel). Store it under a unique constraint and insert with "ignore on conflict":

INSERT INTO notifications (dedupe_key, user_id, …)
VALUES ('u88:evt991:email', 88, …)
ON CONFLICT (dedupe_key) DO NOTHING;

The first insert creates the row. Every replay hits the constraint and does nothing.

Instead of a unique constraint, the planner first checks a Redis set of processed event ids, then inserts. Two planners get the same event at once. What can happen?

Both check, both see it's new, and both insert notifications.

The check and the insert are separate steps. Between them, the other planner can do the same check. Only an atomic operation, like the database's unique constraint, closes that gap.

Why not deduplicate by skipping notifications whose text matches one sent to the same user in the last 5 minutes?

Text isn't identity. Two different people commenting "+1" produce identical text and are genuinely two notifications, so one gets wrongly suppressed. And a replay that arrives six minutes later has a different time window, so it gets through.

Identity has to come from the cause (the event id), not from how the result looks.

What the stage asks

How should the planner avoid creating duplicate notifications?

  1. Sound

    Insert each notification with a unique key of (user, event id, channel); conflicts are ignored

    The identity of a notification is derived from what caused it, so any number of replays converge on the same rows. The database's unique constraint does the deduplication atomically, even with concurrent planners.

  2. Defensible

    Record processed event ids in Redis with a 24-hour TTL and skip events already seen

    It catches most replays, but the check and the insert are separate steps (two planners can both pass the check), the record can be lost on a Redis failover, and replays after 24 hours slip through. Useful as an optimization, not as the guarantee.

  3. Flawed

    Configure the event stream for exactly-once delivery

    Exactly-once features cover the stream's own bookkeeping, not your database writes. A planner that inserts notifications and crashes before committing its offset will see the event again.

  4. Defensible

    Skip a notification if one with identical text was sent to the user in the last 5 minutes

    It suppresses some duplicates, but two genuinely different events can render identical text (two people commenting "+1"), and a replay six minutes later gets through. Similarity is a spam-control tool, not an identity.

What a strong answer covers

  • The notification's identity is derived from its cause (event id), recipient and channel, so replays produce the same identity.
  • Uniqueness is enforced atomically (unique constraint), not by a separate check.
  • The stream's delivery guarantee cannot cover side effects in your database.Supporting

The reasoning

  1. Derive each notification's identity from its cause: recipient, event id and channel.
  2. Enforce uniqueness atomically with a unique constraint; a separate check-then-insert races.
  3. This deduplicates notifications, not sends; a sender can still duplicate later.

A derived key is the most robust form of Idempotency: nobody has to remember an id, because the id is the cause. INSERT … ON CONFLICT (dedupe_key) DO NOTHING makes replays free.

This deduplicates notifications. It does not yet deduplicate sends, because a sender can still crash between calling the provider and recording the result. That gap is where duplicates actually come from, and stage 6 deals with it.

Stage 4 of 11 · Decide

When are preferences applied?

A user turns off comment emails. Notifications already planned for them are sitting in a queue behind an announcement and may not be sent for an hour. The requirement says preference changes take effect immediately.

What you need to know first

A notification is planned at one moment and sent at another. Normally the gap is seconds. Behind a large announcement, or during a provider outage, it can be an hour.

Anything decided at planning time can be out of date by sending time. The question is which decisions must be checked again just before the irreversible step.

A user turns off comment emails at 10:00. Three comment emails for them were planned at 09:55 and are queued behind an announcement until 10:40. If preferences are checked only at planning time, what happens?

All three are sent at 10:40, forty minutes after the user said no. From their point of view, the setting didn't work.

Checking preferences again right before sending would have suppressed all three.

There are good reasons to check at planning time as well. At fan-out scale, not creating unwanted deliveries saves work in every queue and sender.

When a send-time check cancels a delivery, record it as suppressed rather than deleting it. Support can then answer "why didn't I get this?" from the record.

Senders read preferences on every send. To save database load, should they cache them for an hour?

No: an hour-old cache means an unsubscribe can be ignored for an hour. Cache for seconds, or invalidate on change.

The requirement is that changes take effect immediately. A short TTL or explicit invalidation keeps reads cheap without making preferences stale.

What the stage asks

Where should preferences be enforced?

  1. Defensible

    Only when planning: deliveries are created for channels the user allowed at that moment

    It avoids planning unwanted work, which matters at fan-out scale. But anything already queued when the user changes their mind is sent anyway, possibly an hour later, which breaks 'immediately'.

  2. Sound

    Filter when planning, and re-check the current preference just before each provider call

    Planning-time filtering keeps unwanted work out of the queues; the send-time check makes changes effective immediately for anything still queued. Deliveries cancelled at send time get a terminal suppressed state, so the history still explains what happened.

  3. Flawed

    Senders cache preferences for an hour to avoid database load

    An unsubscribe can be ignored for up to an hour, possibly a legal problem for marketing email. If preference reads are expensive, cache with explicit invalidation on change, or keep the TTL in seconds.

  4. Defensible

    Rely on the email provider's suppression list for unsubscribes

    You need the provider's suppression list for legal unsubscribe links and bounces anyway. But it knows nothing about per-type preferences, quiet hours or push, so it is a backstop, not the policy engine.

What a strong answer covers

  • Recognizes the gap between planning and sending, during which preferences can change.
  • A check at send time is required for changes to take effect immediately.
  • Suppressed deliveries are recorded as such, not silently dropped.Supporting

The reasoning

  1. Decisions made at planning time can be stale at sending time; re-check right before the irreversible step.
  2. Filter at planning too, to avoid creating unwanted work at fan-out scale.
  3. Record suppressed deliveries so the history explains what happened.

Any decision made at one time and acted on later can be stale by the time it is acted on. The fix is to re-validate at the point of action, the same move as checking a lease token at completion.

Security notifications bypass this check by type, which is a policy rule written down in one place, not an exception buried in a sender.

Stage 5 of 11 · Model

Trace a mention to a phone

Alice mentions Bob in a comment. Put the steps from her click to the push notification on Bob's phone in order.

What you need to know first

A pipeline that crosses several services stays reliable when every hop follows the same rhythm: record durably, act, record the outcome. Each record lets the next step be retried without redoing or losing work.

When tracing a request, look for where the user's part ends. Everything after that can be slow or retried without them seeing it.

When can Alice's comment request return?

As soon as the comment and its event have committed

From then on the notification is owed and durable. Planning, queueing and sending happen afterwards and don't affect her request.

Push providers (APNs, FCM) report a device token as invalid when the app has been uninstalled or the token has rotated. Sending to it again will fail every time.

A sender that removes tokens as soon as a provider rejects them keeps each user's device list accurate without a separate cleanup job.

Bob's inbox shows the mention even though the push to his phone failed. Why is that possible?

The notification row is created by the planner before any delivery is attempted, and the inbox reads those rows. Push is one delivery of the notification, not the notification itself, so a failed push doesn't affect the inbox.

What the stage asks

Order the life of a mention notification.

In this order

  1. 1Comment and mention event commit in one transaction; Alice's request returns
  2. 2Outbox relay publishes the event to the stream
  3. 3Planner resolves Bob, applies his preferences and quiet hours
  4. 4Planner inserts the push and inbox notifications under dedupe keys
  5. 5Planner enqueues the push delivery on the transactional queue
  6. 6A push sender claims the delivery and marks it sending
  7. 7Sender re-checks Bob's current preferences
  8. 8Sender calls APNs for each of Bob's devices
  9. 9Sender records the outcome and removes any token APNs reports invalid
  • Alice's request ends at the first step. Everything after it can fail and retry without her knowing.
  • Notifications exist (and appear in Bob's inbox) before any delivery is attempted. The inbox does not depend on push succeeding.
  • The preference check happens twice: once to avoid planning unwanted work, and once at the last moment before the irreversible step.
  • Invalid tokens are cleaned up as a side effect of sending, which keeps the device list honest without a separate process.

The reasoning

  1. The user's request ends when the action and its event commit; everything after is asynchronous.
  2. At every hop: record durably, act, record the outcome, so each step can be retried safely.
  3. Notifications exist before delivery, so the inbox doesn't depend on push or email succeeding.

Notice the pattern repeated at every hop: record durably, then act, then record the outcome. The comment records the event; the planner records notifications; the sender records the delivery state. Each record lets the next step be retried safely, and lets the inbox and support tools answer "what happened to this notification?"

Stage 6 of 11 · Break it

Why users got two emails

The notification row is unique. The duplicate happened later, in sending. Find the design decisions that produced it.

What you need to know first

A queue's visibility timeout is a lease: once a sender receives a message, the queue hides it for that long. If the sender doesn't acknowledge in time, the message reappears and another sender can receive it.

So the visibility timeout has to be longer than the slowest the work can take, or the sender has to extend it while it works.

The visibility timeout is 5 s. The provider call's own timeout is 30 s. What can happen?

A slow call is still running when the message reappears, and a second sender sends the same email.

The queue gives up on the first sender after 5 s, while that sender is willing to wait 30 s. For 25 s both can be working on the same delivery.

A second safeguard: record state before the side effect. Before calling the provider, the sender changes the delivery from queued to sending with a fresh attempt token, in one conditional update:

UPDATE deliveries SET status = 'sending', attempt_token = $2
WHERE id = $1 AND status = 'queued'

If another sender already claimed it, zero rows change and this one skips the delivery.

With a long enough lease and the claim, can a duplicate email still happen?

Yes, rarely. If the provider call times out, the sender doesn't know whether the email went out. Retrying may send it twice. Not retrying may mean it was never sent.

The claim prevents two senders working at once. It can't remove the uncertainty of an unconfirmed send without the provider's help.

What the stage asks

Select the lines where the design is at fault.

Logdelivery d-5521 (email, user 88)
  1. 109:00:00.000 sender-2 receives d-5521 from queue (visibility timeout 5 s)
  2. 209:00:00.010 sender-2 calls email provider (client timeout 30 s)

    The visibility timeout (5 s) is shorter than the provider call can take (30 s). The message becomes visible again while the first send is still in flight. Extend the visibility timeout while working, or make it longer than the call's own timeout.

  3. 309:00:05.000 d-5521 visible again (no ack yet)
  4. 409:00:05.040 sender-7 receives d-5521 and calls email provider

    The second sender does not check the delivery's state. A 'sending' state recorded before the call (with an attempt token) would show another attempt is in progress.

  5. 509:00:06.900 provider accepts sender-2's request: msg-a
  6. 609:00:07.300 provider accepts sender-7's request: msg-b
  7. 709:00:07.310 sender-2 marks d-5521 sent; sender-7 marks d-5521 sent

    Both completions succeed because the update is unconditional. Conditioning it on the attempt token would at least detect, and record, the duplicate.

What the fix has to do

  • The queue lease expired while the send was still in flight, so a second sender received the same delivery.
  • Recording a 'sending' state with an attempt token before calling the provider lets other senders see and skip in-flight work.
  • Without provider idempotency, a send whose outcome is unknown can still produce a duplicate on retry: the goal is to make that rare, not impossible.

The reasoning

  1. A queue lease shorter than the work lets a second worker take the same message mid-flight.
  2. Record 'sending' with an attempt token before the side effect, so concurrent senders skip in-flight work.
  3. Without provider idempotency, unknown outcomes can still duplicate; make that rare and choose deliberately.

Two separate ideas apply. Leases must outlive the work they protect, or be extended by heartbeats; this is the video pipeline's lesson again. State before side effect: mark the delivery sending with an attempt token, and treat an existing sending that is recent as "someone else has it".

What remains is honest uncertainty. If a send times out, the email may or may not have gone out. You choose: retry (risking a duplicate) or not (risking silence). For a security alert, retry; for a comment notification, perhaps not. That choice is a product decision, not a technical one. See Delivery guarantees and Leases and fencing tokens.

Stage 7 of 11 · Break it

The email provider is failing

The requirement: outages delay notifications but never lose them. Nothing about the outage is under your control except how you respond to it.

What you need to know first

Per-message retries with backoff assume failures are independent: this message failed, the next one might not. A provider outage breaks that assumption. Every message fails for the same reason at the same time.

A circuit breaker watches failures across all calls to a dependency. When the failure rate crosses a threshold it "opens": calls stop for a while. Then it lets a few probe calls through, and closes again once they succeed.

The email queue grows by 400 messages a second during a 25-minute outage. About how many messages are waiting when it ends?

About 600,000 messages.

400 × 25 × 60 = 600,000 messages. At the 500-a-second limit, draining them takes 1,200 s, about 20 minutes, on top of normal traffic. That's why the recovery drain needs the same rate limiting and priorities as normal sending.

Each message retries with exponential backoff and is marked failed after 5 attempts over about 2 minutes. The outage lasts 25 minutes. What happens?

Messages exhaust their attempts during the outage and are marked failed: notifications are lost.

The retry budget is shorter than the outage. Per-message policies can't tell 'this message is bad' from 'the provider is down'.

Failures also differ in kind, and each kind deserves a different response:

ResponseKindWhat to do
429, 503transientretry later, with backoff
invalid address, unsubscribedpermanentstop; mark failed
timeoutunknowna policy choice: retry and risk a duplicate, or don't

What the stage asks

How should senders behave during the outage?

  1. Flawed

    Retry each failed send immediately, up to 10 times

    Every failure multiplies into ten requests against a provider that is already overloaded and rate-limiting you, which prolongs the outage and burns through attempts that could have succeeded later.

  2. Sound

    Trip a circuit breaker for the email channel: pause sending, probe periodically, keep work queued, then resume within the rate limit

    Senders stop hammering a failing provider and let the queue absorb the outage. Deliveries keep their place and their state, so nothing is lost. On recovery, the shared rate limiter meters the backlog out at 500/s, critical class first.

  3. Defensible

    Retry with exponential backoff per message, and mark it failed after 5 attempts

    Backoff is right for an individual message's transient errors. But during a 25-minute outage every message burns its attempts and fails, so notifications are lost, which the requirement forbids. Channel-wide outages need channel-wide handling.

  4. Defensible

    Fail over to a second email provider immediately

    A real option, which many large senders keep warm. But sends whose outcome was unknown at the first provider may be duplicated at the second, and a cold sending domain can land in spam. It works as a planned capability, not a reflex.

What a strong answer covers

  • Treats a provider-wide outage at the channel level (circuit breaker) rather than per message.
  • Avoids amplifying load on a failing dependency (backoff, pause, probe).
  • Work stays durably queued, so the backlog drains on recovery.
  • The recovery drain respects the rate limit and priority classes.Supporting

The reasoning

  1. Provider-wide outages need a channel-level response: pause, probe, then resume.
  2. Keep deliveries durably queued during the outage; per-message retry budgets would otherwise lose them.
  3. Classify failures as transient, permanent or unknown, and drain the backlog within the rate limit.

Per-message retry policies assume failures are independent. Provider outages are the opposite: every message fails for the same reason at the same time. The right unit of response is the channel. Stop, wait, probe, then drain at a controlled rate; see Retries, backoff and jitter and Backpressure and capacity.

Classify errors too: a 429 or 503 is transient; "invalid address" or "unsubscribed" are permanent and terminal; a timeout is unknown.

Stage 8 of 11 · Change it

The announcement and the security alert

Same provider, same account, same 500/s.

What you need to know first

Two different problems hide behind "urgent messages wait":

  • Order: urgent work is behind other work in the same queue. Priorities or separate queues fix this.
  • Capacity: even at the front of the queue, there is no capacity left to serve it. Only reserving capacity fixes this.

Here the capacity is the provider's 500 emails a second for the whole account.

Alerts get the highest priority in a shared queue, but bulk senders have already used this second's 500 sends. When does the alert go out?

When quota frees up, competing with bulk senders that are already waiting for the same quota

Priority decides which message a sender picks up, but every sender is waiting for the same exhausted quota. The alert has no capacity of its own.

Reserving capacity means giving each class its own token bucket that draws from the account limit, for example critical 50/s, transactional 150/s, bulk 300/s. Bulk can't take what's reserved for critical, so a critical alert waits only for other critical alerts.

Unused reserved tokens can be lent to bulk each second, so reservation costs little when there are no alerts.

With 300 of the 500 sends a second reserved for bulk, about how many hours does the 5-million-user announcement take?

About 4.6 hours.

5,000,000 ÷ 300 ≈ 16,700 s ≈ 4.6 hours, compared with 2.8 hours at the full 500. That is the cost of guaranteeing the alerts, and less in practice when reserved capacity is lent to bulk while idle.

What the stage asks

How do you keep security alerts fast during the announcement?

  1. Flawed

    One queue in arrival order; the announcement goes first because it was scheduled first

    Alerts wait behind millions of messages for hours.

  2. Defensible

    One queue with a priority field; senders pick the highest priority first

    Ordering improves, but quota is the bottleneck: if bulk sends have already used this second's 500 tokens, the alert still waits, and many queues cannot do priority efficiently at millions of messages. Priority changes order; it does not reserve capacity.

  3. Sound

    Separate queues per class (critical, transactional, bulk) with a reserved share of the provider quota for critical and transactional

    Bulk can never consume the capacity reserved for critical, so an alert waits at most for its own class's small queue. Unused reserved capacity can be lent to bulk each second. The cost is a slightly slower announcement.

  4. Sound

    Send marketing through a separate provider account or subdomain

    Physically separate quotas and sender reputation: a spam complaint about marketing cannot hurt delivery of password-reset emails. Many teams do this alongside class-based queues.

What a strong answer covers

  • The constraint is shared capacity, so the fix must reserve capacity, not only reorder work.
  • Bulk and critical traffic are isolated in queueing and quota (or accounts).
  • Names the cost: slower bulk sends, unused reserved capacity, or extra accounts.Supporting

The reasoning

  1. A latency guarantee under saturation needs reserved capacity, not just priority.
  2. Give each class its own token bucket drawing from the shared provider limit, and lend unused capacity to bulk.
  3. Separate accounts or domains isolate sender reputation as well as quota.

This is the fast-lane problem from the video pipeline in a different system: a latency guarantee under saturation needs reserved capacity. Here the capacity is a provider's rate limit rather than workers, and the mechanism is a token bucket per class drawing from the account's 500/s.

Separate accounts add another kind of isolation, of reputation: bulk email gets spam complaints, and you do not want those to affect password resets.

Stage 9 of 11 · Change it

Digests and quiet hours

Batching changes notifications from "send when it happens" to "send what accumulated". Evaluate these statements about the new behaviour.

What you need to know first

A digest replaces "send when it happens" with "periodically, send what accumulated". Two new things need care:

  • The digest itself is a notification, sent by a scheduled job that can run twice, so it needs an identity too, such as (user, issue, hour window).
  • Membership: which notifications went into which digest. If that isn't recorded atomically with the digest, a crash can put one notification in two digests, or in none.

A digest job creates the digest record, crashes, and never marks the included notifications as 'in a digest'. The next run starts. What happens?

Those notifications still look un-digested, so the next run includes them again. If the first digest was sent, the user gets them twice; if it wasn't, the first digest is an orphan record.

Creating the digest and marking its members in one transaction removes both outcomes.

"Quiet hours from 22:00 to 07:00" means the user's 22:00. A server running in UTC has to convert using the user's time zone, including daylight-saving changes, which shift the offset twice a year in many regions.

A user in Tokyo (UTC+9) set quiet hours 22:00–07:00. The server evaluates them in UTC. At 23:00 Tokyo time, is a notification held?

No: 23:00 in Tokyo is 14:00 UTC, outside 22:00–07:00 UTC, so it's sent in the middle of their night.

Evaluating in server time shifts quiet hours by the user's offset. They need to be checked in the user's own time zone.

What the stage asks

Which statements hold?

  1. Holds

    A digest job that runs twice for the same user and hour must not send two digests.

    Scheduled jobs are retried and overlap like everything else. Give the digest a derived identity, such as (user, issue, hour window), with a unique constraint, exactly like individual notifications.

  2. Holds

    Notifications included in a digest must be marked as included in the same transaction that creates the digest.

    Otherwise a crash between the two either sends a notification twice (in two digests) or never (marked but the digest is lost). Membership and the digest record commit together; sending follows.

  3. Fails

    Quiet hours can be evaluated in the server's time zone.

    Quiet hours are the user's night, not the data centre's. Store a time zone per user and evaluate their local time, including daylight-saving transitions.

  4. Fails

    Security alerts should be held until a user's quiet hours end.

    A new-login alert at 3 a.m. may be the only chance to stop an account takeover. Policy rules by type, such as 'critical ignores quiet hours', must be explicit.

  5. Depends

    A longer digest window always produces a better user experience.

    Fewer emails, but each one is later. For a busy issue, hourly batching is a relief; for an assignment with a deadline, an hour can be too long. That is why digests are usually per type and user-configurable.

The reasoning

  1. A digest is a notification too: give it a derived identity so a re-run can't send it twice.
  2. Record digest membership in the same transaction that creates the digest.
  3. Quiet hours are the user's local time; critical types bypass them by explicit policy.

Digests introduce aggregation, and aggregation needs the same care as everything else: a derived identity for the aggregate, atomic membership, and the side effect only after the record. The time-related claims are reminders that "when" is a user-facing concept: quiet hours and digest windows live in each user's local time.

Stage 10 of 11 · Change it

Write the sender

Write the core loop of an email sender: claim a delivery, check it is still wanted, respect the class's quota, send, and record the outcome. Handle transient, permanent and unknown outcomes differently.

What you need to know first

A delivery moves through a small set of states. Each move is a conditional update: it only happens if the row is in the expected state, so concurrent senders can't both make it.

FromToWhen
queuedsendinga sender claims it (with a fresh attempt token)
sendingsentthe provider accepted it
sendingsuppressedpreferences no longer allow it
sendingfaileda permanent error, or a deliberate give-up
sendingqueueda transient error; try later

Why condition the final update on the attempt token, not just the delivery id?

So a slow, stale attempt can't overwrite the outcome recorded by a newer attempt

If an attempt was presumed dead and the delivery was re-queued and claimed again, the first attempt may still finish later. Its token no longer matches, so its update changes nothing.

The order of steps inside one attempt matters:

  1. Claim (queued → sending).
  2. Re-check preferences: the last chance to cancel before something irreversible.
  3. Take a token from the class's rate limit.
  4. Call the provider.
  5. Record the outcome, conditioned on the attempt token.

A sender claims a delivery and the machine dies mid-call. The delivery stays in 'sending' forever. What needs to exist?

A reaper: a periodic job that finds deliveries in sending for far longer than any send could take (say 10 minutes), and moves them back to queued for another attempt. Without it, a crashed sender silently loses work.

What the stage asks

Implement processDelivery for the email channel.

Reference implementation

async function processDelivery(deliveryId: string) {
  const token = crypto.randomUUID();
  const d = await db.oneOrNone(`
    UPDATE deliveries SET status = 'sending', attempt = attempt + 1,
           attempt_token = $2, updated_at = now()
    WHERE id = $1 AND status = 'queued' AND next_attempt_at <= now()
    RETURNING *`, [deliveryId, token]);
  if (!d) return; // someone else has it, or it is finished

  const n = await notifications.load(d.notification_id);
  if (!(await policy.allows(n.user_id, n.type, "email"))) {
    return finish(d, token, "suppressed", "preference changed");
  }

  await quotas.take(`email:${d.class}`); // waits for a token from the class's bucket

  try {
    await email.send(render(n), { timeoutMs: 10_000 });
    return finish(d, token, "sent");
  } catch (err) {
    if (isPermanent(err)) return finish(d, token, "failed", err.code); // bad address, unsubscribed
    if (isTimeout(err) && d.class !== "critical") {
      // Unknown outcome: for non-critical mail, prefer silence to a likely duplicate.
      return finish(d, token, "failed", "unknown outcome after timeout");
    }
    // Transient (or critical and unknown): try again later.
    const delay = backoffWithJitter(d.attempt);
    await db.none(`
      UPDATE deliveries SET status = 'queued', next_attempt_at = now() + $3 * interval '1 second',
             last_error = $4
      WHERE id = $1 AND attempt_token = $2`, [d.id, token, delay, String(err)]);
  }
}

async function finish(d: Delivery, token: string, status: string, reason?: string) {
  await db.none(`
    UPDATE deliveries SET status = $3, last_error = $4
    WHERE id = $1 AND attempt_token = $2`, [d.id, token, status, reason ?? null]);
}
  • The claim is a conditional update: queued → sending with a fresh token. Concurrent senders and redelivered queue messages fall through harmlessly.
  • The preference check is the last thing before the irreversible step.
  • Unknown outcomes get a deliberate policy: critical mail is retried (a duplicate password-change alert beats a missing one); routine mail is not (a missing comment email beats a duplicate). Making that choice explicit in code is the point.
  • A reaper re-queues deliveries stuck in sending far longer than any send could take, which covers senders that crashed mid-call.

What a strong answer covers

  • Atomically moves the delivery to 'sending' with a fresh attempt token, and skips it if another attempt is in progress or it is no longer queued.
  • Re-checks current preferences (and type policy) before sending, recording 'suppressed' if no longer wanted.
  • Takes a token from the class's rate limit before calling the provider.
  • Treats transient errors (retry later with backoff), permanent errors (terminal failure) and timeouts (unknown) differently.
  • Records the outcome conditioned on its attempt token.Supporting

The reasoning

  1. Model each delivery as a state machine and make every transition a conditional update.
  2. Re-check preferences and take quota immediately before the irreversible provider call.
  3. Treat transient, permanent and unknown failures differently, and reap deliveries stuck in 'sending'.

The sender is a small state machine with a policy at each transition: whether a send is still wanted, whether quota is available, and what kind of failure occurred. Most of the sender's correctness lives in the WHERE clauses, which is where it lives in every system in this course.

Stage 11 of 11 · Defend it

Defend the duplicate policy

The product manager asks for a simple promise in the help centre: "We never send the same notification twice." Explain what you can promise, what you cannot, and why.

What you need to know first

Explaining a guarantee to a non-engineer comes down to three plain statements:

  • What we guarantee, in terms of what the person sees.
  • What we can't, and the one concrete situation where it happens.
  • What we chose in that situation, and why it's the better failure for them.

Avoid words like "idempotent" or "at least once". Describe the outcome.

Which public promise stays true during incidents?

"Duplicates are rare and can only happen when an email provider fails mid-send. Security alerts are always retried."

It names the exact residual case and the deliberate choice. It is true on the worst day, not just the average one.

What the stage asks

Explain to a non-engineer what the system guarantees about duplicates and missed notifications, and why a stronger promise is not possible.

Reference answer

What we guarantee: each event produces at most one notification per person per channel, however many times our systems process it. Two of our servers never send the same notification at the same time.

What we cannot guarantee: when the email provider does not confirm a send (a timeout, a dropped connection), we cannot know whether the email went out. The provider offers no way to ask "did you already send this?". At that moment there are only two options: send again and risk a duplicate, or do not and risk the person never getting it.

What we chose: for security alerts we resend, because a duplicate is better than a missing warning. For everything else we do not, because a missing comment notification is better than an annoying duplicate, and the inbox still shows it.

The honest promise: "Duplicate notifications are rare; they can happen only when an email provider fails in the middle of sending. Security alerts are always retried." It is less catchy than "never", but it stays true during incidents.

What a strong answer covers

  • States what is guaranteed: one notification per event per channel (dedupe keys), and no concurrent duplicate sends (claims).
  • Explains the residual case plainly: when the email provider does not confirm a send, we cannot know whether it was delivered.
  • Explains the deliberate choice per type: retry critical alerts (risking a duplicate), do not retry routine ones (risking a miss).
  • Proposes an honest public promise, e.g. 'duplicates are rare and only occur when a provider fails mid-send'.Supporting

The reasoning

  1. 'Never' is a claim about every failure mode, including ones in systems you don't control.
  2. Explain guarantees as outcomes: what's guaranteed, the one case that isn't, and the choice made there.
  3. Prefer a slightly weaker promise that stays true during incidents.

The point is that "never" is a claim about every failure mode, including ones in systems you do not control. Engineers earn trust by making promises they can keep and explaining the trade they chose in terms the business can weigh. That is the same skill as defending a design in an interview, aimed at a different audience.

How you did

Now try it as an interview question

  • “Design a notification system.”
  • “How would you send an email to every user without affecting transactional email?”
  • “Users are getting duplicate push notifications. How do you investigate and fix it?”
  • “Design a system that respects user notification preferences at scale.”

The interview mode mixes stages from this and other investigations with concept recall and questions about your own projects.

Back to the last stage