Stage 1 of 11 · Model
What the numbers say
Start with arithmetic and the provider's contract. Eight million notifications a day, a 500/s email limit, a five-million-user announcement, and a 30-second promise for security alerts.
What you need to know first
When a provider caps your sending rate, that cap is a shared budget. Every email of every kind draws from it: security alerts, comment notifications and marketing alike. The first numbers to work out are how long the big jobs occupy that budget, and what is left for everything else meanwhile.
An announcement goes to 5 million users, and the provider allows 500 emails a second. About how many hours does it take?
About 2.8 hours.
5,000,000 ÷ 500 = 10,000 seconds ≈ 2.8 hours. For that whole time, the account has no spare capacity unless something reserves it.
8 million notifications a day, with work-hour peaks at 10× the average. About how many a second at peak?
About 930 per second.
8,000,000 ÷ 86,400 ≈ 93 a second on average. × 10 ≈ 930 a second at peak.
If most of them are emails, the peak alone is nearly twice the provider's 500 a second. Something has to wait, and the design decides what.
An idempotency key is an id you send with a request so the receiver can recognise a retry: "you already did 812, here's the same result". Payment APIs usually accept one. This email provider doesn't.
Without it, a request that times out is ambiguous: the email may or may not have been sent, and the provider gives you no way to ask.
A send to the provider times out after 10 seconds. What do you know?
Nothing for certain: it may have been sent, or not.
The timeout is on your side. The provider may have sent the email and its response got lost, or it may never have received the request. Retrying risks a duplicate; not retrying risks silence.
What the stage asks
Which statements hold?
- Holds
Sending a marketing email to all 5 million users takes nearly three hours at the provider's limit.
5,000,000 / 500 per second = 10,000 seconds, about 2.8 hours. For that long, the account's entire send capacity is in use.
- Holds
If security alerts share a FIFO queue with the announcement, they can wait hours.
A security alert enqueued just after the announcement waits behind millions of messages draining at 500/s. The 30-second promise requires isolation from bulk traffic, both in queueing and in quota.
- Fails
Sending notifications inside the request that created the comment is fine, since each comment notifies only a few people.
The count is small but the dependency is the problem: provider latency becomes comment latency, and provider outages become comment failures, which is last week's incident. The product action must succeed without the notification having been sent.
- Fails
With careful engineering, every email can be guaranteed to be delivered exactly once.
The provider accepts no idempotency key. If a send times out, you cannot know whether it went out; retrying risks a duplicate and not retrying risks silence. You can make duplicates rare, but not impossible. See Delivery guarantees.
- Fails
Eight million notifications a day is about 93 a second, so peak capacity is not a concern.
Averages hide peaks. Work hours run around 10x the average, roughly 900 a second, which already exceeds the email limit if most notifications are email. Size for peaks, and decide what waits.
The reasoning
- A provider's rate limit is a shared budget; bulk sends can occupy all of it for hours.
- Size for peaks, not averages, and decide in advance what waits when demand exceeds the limit.
- Without provider idempotency keys, a timed-out send is ambiguous, so exactly-once email is impossible.
The provider's limit makes email capacity a shared, scarce resource, and the announcement can consume all of it for hours. Most of the design is about who gets that capacity and when.
The missing idempotency key is the other defining fact: for email, "exactly once" is off the table, so the goal becomes at most once per notification, with the residual duplicate window as small as possible.