Skip to content

The finished design, decision by decision

How to design a Payment System

Not the one correct diagram, but a design you can defend under these constraints: the finished architecture, then every stage's question with the reasoning that answers it, the tradeoffs it accepts, and where another engineer could land differently.

A payment workflow that never double-charges

The short answer

8 parts, each with one job. The map below shows how requests and data move between them; the stages after it explain why each part is there.

Buyer's browser
Collects card details in the provider's hosted fields, completes 3-D Secure, and submits checkout with an idempotency key; polls for status.
Checkout API
Claims the attempt by idempotency key, calls the provider with the same key, and applies results as conditional transitions.
Postgres
System of record for orders, payment attempts, the event ledger and the outbox.
Payment provider
Moves the money. Source of truth for every charge and refund.
Webhook receiver
Verifies signatures, records the event id, and applies forward-only transitions; acknowledges within a second.
Fulfillment worker
Processes outbox rows: grants enrollment, revokes it on refund, sends receipts. Every handler is idempotent.
Email provider
Delivers receipts.
Reconciler
Resolves attempts stuck in processing by asking the provider, and diffs the daily settlement report against the ledger.
12345678910CLIENTBuyer's browserSERVICECheckout APIDATABASEPostgresEXTERNALPayment providerSERVICEWebhook receiverWORKERFulfillmentworkerEXTERNALEmail providerWORKERReconciler

Select a component to see what it is responsible for and which state it owns.

  1. 1Buyer's browser → Payment provider: Card entry and 3-D Secure in hosted fields
  2. 2Buyer's browser → Checkout API: Pay (Idempotency-Key header); poll status
  3. 3Checkout API → Payment provider: Create payment with the attempt's key
  4. 4Checkout API → Postgres: Claim attempt; conditional transitions
  5. 5Payment provider → Webhook receiver: Signed events: at least once, unordered
  6. 6Webhook receiver → Postgres: Dedupe by event id; forward-only transition
  7. 7Fulfillment worker → Postgres: Claim outbox rows; record completion
  8. 8Fulfillment worker → Email provider: Send receipt with idempotency key
  9. 9Reconciler → Postgres: Attempts processing too long; ledger
  10. 10Reconciler → Payment provider: Look up by reference; settlement report
  • Request / response
  • Asynchronous

Why does this design work?

The design never lets a request be the unit of reliability. Requests are duplicated, cut off and abandoned. The unit is the attempt: a durable row created atomically under an idempotency key before anything irreversible happens, carrying that key into every call that could move money.

From there, every source of information (the synchronous response, webhooks, reconciliation) feeds the same forward-only, conditional transitions. They can arrive in any order, any number of times, and the state still converges on what the provider knows. Side effects are recorded in the same transaction as the transitions that cause them, so they happen at least once and are deduplicated where they land.

Uncertainty is a first-class state. The system never guesses whether a timed-out charge succeeded; it waits to be told, and reconciliation guarantees that it eventually is.

Invariants, and where they are enforced

  • At most one charge exists for each payment attempt.

    The attempt row is claimed with a unique idempotency key before any provider call, and every call for that attempt carries the same key, so both your database and the provider deduplicate. Enforced by Postgres, Checkout API, Payment provider.

  • An order is paid only if the provider reported the payment as succeeded.

    Only provider-originated facts (synchronous response, verified webhook, reconciliation lookup) can drive processing → succeeded, through the same conditional update. Enforced by Checkout API, Webhook receiver, Reconciler.

  • Terminal states never move backwards.

    Every transition is an UPDATE conditioned on the allowed previous states; late or duplicate events match zero rows. Enforced by Postgres.

  • Enrollment and receipts happen if and only if the payment succeeded, eventually and without duplicates.

    The transition and its outbox rows commit in one transaction; fulfillment handlers are idempotent (unique enrollment per user and course, email keyed by attempt). Enforced by Postgres, Fulfillment worker.

What does it rely on?

  • The provider honours idempotency keys for 24 hours and allows lookup by your reference.
  • Clients reuse the same key when retrying the same attempt.
  • All money movement goes through the attempt table; dashboard actions are caught only by reconciliation.
  • Postgres is available for claims and transitions. When it is down, checkout fails closed rather than charging unrecorded.
  • Fulfillment handlers are idempotent.

What tradeoffs does it make?

ChoiceGainsCosts
Bounded synchronous wait plus a processing stateImmediate answers for most buyers; correctness for the rest.Two resolution paths and a status page to build.
Attempts plus an append-only ledgerPer-attempt idempotency, auditability, room for refunds and retries.More schema, and derived order status.
Transactional outbox for side effectsNo lost enrollments or receipts after crashes.A worker to run, and idempotency required in every handler.
Reconciliation as a permanent componentCertainty that every attempt resolves; early detection of broken paths.Ongoing API usage, and a human queue for true anomalies.

What are the reasonable alternatives?

Fully asynchronous charging through a job queue
Better when charges are off-session (renewals, payouts) and nobody is waiting at a checkout page.
Provider-hosted checkout pages with webhook-only integration
Better when you want the smallest possible payment surface: the provider owns the payment UI and retries, and you own only webhook handling and reconciliation.
Workflow engine (Temporal, Step Functions) orchestrating each payment
Better when payments involve many steps, timers and human approvals, such as marketplaces with escrow and payouts.

When does it stop working?

  • Clients generate a new idempotency key for every retry, turning retries into new attempts.
  • Engineers fail over to another provider, or mark attempts failed, while outcomes are still unknown.
  • Money moves outside the system (dashboard refunds, scripts) faster than reconciliation catches it.
  • Provider idempotency semantics differ from the assumed contract, for example keys scoped per API version or retained for a shorter time.

Every stage, decided and explained

Spoilers, for the whole investigation: each stage's question and its answer, the reasoning behind it, and the tradeoffs it accepts. If you have not worked through the stages yet, you may want to do that first.

Work through the stages

Stage 1 of 13 · Model

What the provider's behaviour implies

The provider's documentation is a list of facts. A design starts by turning each one into a consequence. Before you write any code, decide which of these statements follow from the constraints.

What you need to know first

A timeout ends your waiting, not the other side's work. When a request to the provider times out, three things are possible: the request never arrived, it arrived and failed, or it arrived and succeeded and only the response was lost.

From your side these look identical. A design that treats a timeout as "failed" will one day mark a successful charge as failed and let the buyer pay again.

An idempotency key is a value you send with a request so the receiver can recognise a repeat. This provider stores each key for 24 hours: a second request with the same key returns the first result instead of charging again. See Idempotency.

A webhook is the provider calling your server when something changes. These are delivered at least once (possibly twice), in no guaranteed order, and retried for three days if your endpoint fails.

A request with idempotency key K succeeds. 30 hours later, a buggy client retries with the same key K. What does the provider do?

Creates a new charge: it has forgotten K.

The provider only remembers keys for 24 hours. After that, the same key looks new. Your own records have to catch late retries.

At peak, 50 orders a minute. About how many requests a second is that to the provider, counting one create per order?

About 0.83 per second.

50 ÷ 60 ≈ 0.83 a second. The provider allows 100. Scale isn't the problem in this design; uncertainty and concurrency are.

What the stage asks

Which statements follow from how the provider behaves?

  1. Fails

    If the call to create a payment times out, the buyer was not charged.

    The request may have reached the provider and succeeded, with only the response lost or late. A timeout ends your waiting, not the provider's work. The outcome is unknown; see Timeouts and unknown outcomes.

  2. Holds

    Repeating a create-payment request with the same idempotency key within 24 hours returns the original result instead of charging again.

    That is the contract the provider offers, and the foundation of the whole design. It also implies that a retry after 24 hours, or with a new key, is a brand-new charge.

  3. Fails

    Because webhooks are retried for three days, they are a complete record of every outcome.

    If your endpoint is broken for longer than the retry window, misconfigured, or rejecting signatures after a secret rotation, events are lost for good. And until one arrives, you cannot tell "not yet" from "never". Something has to ask the provider; see Reconciliation.

  4. Holds

    A payment.processing webhook can arrive after the payment.succeeded webhook for the same payment.

    Each event is delivered and retried independently. If the first delivery of processing failed, its retry can land after succeeded. Handlers must never let an older event regress newer state.

  5. Fails

    At 50 orders a minute, the provider's 100 requests/second limit constrains the design.

    50 a minute is under one request per second, two orders of magnitude below the limit, even with retries and lookups. It is worth knowing the limit exists, but it does not shape this design. It will at 100x.

The reasoning

  1. A timeout means the outcome is unknown; the charge may have succeeded.
  2. Idempotency keys deduplicate retries only within the provider's retention window.
  3. Webhooks arrive late, twice, out of order or not at all, so something must ask the provider directly.

Three facts drive everything that follows:

  • Outcomes can be unknown. Any call can end without telling you what happened. The design needs a state for "we asked and do not know yet" and a way to find out.
  • The provider deduplicates by key, for 24 hours. Every retry of the same intent must carry the same key, and the key must be tied to a record you write before calling.
  • Events are hints, not commands. Webhooks arrive late, twice or out of order, and sometimes not at all. They must be applied in a way that tolerates all four.

Notice what is absent: nothing in the brief is a scale problem. The difficulty in payments is almost entirely uncertainty and concurrency, not throughput.

Stage 2 of 13 · Decide

Should checkout wait for the provider?

The buyer clicks Pay. Most charges complete in 1-2 seconds, a few take 10, some hang, and 3-D Secure payments complete only after the buyer finishes a challenge. The load balancer cuts requests at 30 seconds. The buyer should get an answer or a clear "processing" state within about 10 seconds.

What you need to know first

A request has a deadline set by things outside your code: the load balancer here cuts requests at 30 seconds. If your code waits longer than that, the buyer sees the load balancer's generic error, and your handler loses control of the response.

So any wait inside a request needs its own timeout, set shorter than the load balancer's, with a plan for what to return when it fires.

Provider p99 latency is 10 s. With an 8-second wait, roughly what percentage of buyers would see 'processing' instead of an immediate answer?

About 1.5 %.

p99 = 10 s means 1% of charges take longer than 10 s. An 8 s cut-off catches a little more than that: roughly 1 to 2%. The other 98% get their answer in the request.

The "processing" path only works if something will resolve it later. That's possible only if a durable record of the attempt exists before the provider is called, so a webhook or a periodic check can find it and finish it.

3-D Secure payments always take this path: the buyer completes a challenge minutes later, and the result arrives by webhook.

The browser confirms a payment with the provider's SDK. Why shouldn't the server mark the order paid when the browser says so?

The browser's report can be forged or never sent; the server must learn the outcome from the provider.

Anyone can call your API saying 'I paid'. And a tab that closes right after paying never reports at all. The provider's response, a verified webhook or a lookup are the only trustworthy sources.

What the stage asks

How should the checkout request relate to the provider call?

  1. Defensible

    Call the provider inside the request and wait up to 25 seconds for the result

    It works for the common case, and most simple integrations start here. But a hang that reaches the load balancer's timeout leaves the buyer with an error page while the charge may still succeed, and 3-D Secure cannot complete inside a request at all. You still need everything the asynchronous path needs; you have just hidden the need until the first incident.

  2. Sound

    Record the attempt, call the provider with an ~8 s timeout; return the result if it arrives, otherwise return 'processing' and resolve it later

    The common case gets an immediate, honest answer. The uncommon case (slow, hung, or awaiting 3-D Secure) gets an equally honest one: "processing", with the page polling the status. The attempt is already recorded, so webhooks or reconciliation can resolve it whenever the outcome becomes known.

  3. Defensible

    Enqueue a charge job and return 'processing' immediately; a worker calls the provider

    This is robust, and appropriate when charges are not interactive (subscription renewals, batch payouts). For a buyer staring at checkout, it adds a queue hop and a polling delay to a call that usually finishes in 1.5 seconds, and 3-D Secure needs the buyer present anyway.

  4. Flawed

    The browser confirms the payment with the provider's SDK, then tells the API the order is paid

    Confirming in the browser is common and fine. Trusting the browser's report is not. A buyer can forge the call, and a tab that closes after paying never tells you at all. The server must learn the outcome from the provider directly (API response, verified webhook or lookup).

What a strong answer covers

  • Most charges finish in seconds, so a short bounded wait gives most buyers an immediate answer.
  • When the wait ends without a result the outcome is unknown, so the design needs an explicit processing state resolved later by a source of truth.
  • The server learns outcomes from the provider, never from the client's claim.Supporting

The reasoning

  1. Wait synchronously when it's usually fast, with a timeout shorter than the load balancer's.
  2. Record the attempt before the call so a 'processing' result can be resolved later.
  3. Learn outcomes from the provider (response, verified webhook, lookup), never from the client's claim.

This is a hybrid of synchronous and asynchronous: synchronous when it can be, asynchronous when it must be. The request is a fast path layered over a design that would be correct without it.

The 8-second bound is chosen from the requirement ("an answer within about 10 seconds") and the provider's latency (p99 of 10 seconds means about 1-2% of buyers see "processing" briefly). The request timeout must be shorter than the load balancer's, so your code decides what the buyer sees rather than the load balancer's generic 504.

Most importantly, the attempt row exists before the provider call. Whatever happens next, whether a crash, a timeout or a deploy, there is a durable record that a charge may be in flight; see Asynchronous processing.

Tradeoffs

ChoiceGainsCosts
Bounded wait with processing fallbackImmediate answers for ~98% of buyers; correct for the rest.Two code paths for outcomes, and a status page the buyer may have to wait on.

Where another engineer could land differently

For off-session charges such as renewals, the fully asynchronous design is simpler and better: nobody is waiting.

Stage 3 of 13 · Decide

Where does payment state live?

A buyer may try one card, get declined, and pay with another. A payment may sit in "processing" for minutes. Finance needs to know exactly which provider event made an order paid, and support needs to answer "why was I charged?" months later.

What you need to know first

A boolean can say "paid" or "not paid". A payment can be in more states than that: we asked and don't know yet, declined (with a reason), succeeded, refunded. And one order can have several attempts: a declined card, then a different card.

A model with too few states forces you to guess whenever reality is in a state you can't represent.

Payment state is a 'paid' boolean on the order. The provider call times out. What do you store?

You have to guess true or false, and either can be wrong.

There's no value for 'unknown'. False risks the buyer paying twice; true risks granting a course that was never paid for.

Two kinds of table solve different problems:

  • Current state, one row per attempt, updated as it moves: what decisions read.
  • An append-only ledger, one row per change and never updated: who or what changed the state, when, and on which evidence. It's what finance and support read. See Append-only logs.

Writing both in the same transaction means they can never disagree.

A buyer's first card is declined and they pay with a second. Why does each attempt need its own idempotency key?

The provider would treat a reused key as a retry of the first attempt and return the cached decline. The second card would never be tried. A new attempt is a new intent, so it needs a new key; retries of one attempt share its key.

What the stage asks

How should payment state be modelled?

  1. Flawed

    A paid boolean on the orders table

    There is nowhere to represent "we asked and don't know yet", a decline reason, a second attempt, or a refund, and no record of what made it true. The first timeout forces you to guess between true and false, and both guesses can be wrong.

  2. Defensible

    A status enum on the order (pending, paid, failed, refunded), updated in place

    Much better: it is a real state machine. But a declined attempt followed by a new card overwrites the first attempt's history, there is one place for one idempotency key, and "why is this paid?" has no answer beyond the current value. It works for simple stores where each order gets exactly one attempt.

  3. Sound

    A payment_attempts table (one row per attempt, its own key and state machine) plus an append-only payment_events ledger

    Each attempt has its own identity, idempotency key and lifecycle (created → processing → succeeded | failed, succeeded → refunded). The order's status is derived from its attempts, and every transition appends an event recording the cause (API response, webhook evt_…, reconciliation). That makes it auditable and debuggable.

  4. Flawed

    Store only the provider's payment id; ask the provider whenever you need the status

    The provider is the source of truth for money, but not a database for your product. Every order page becomes a remote call subject to rate limits and outages. And you cannot record the attempt before calling the provider, because you have no ID yet, which is exactly when you most need a record.

What a strong answer covers

  • Each attempt (e.g. a declined card, then a new one) needs its own record and its own idempotency key.
  • The model includes an explicit state for outcomes that are not yet known.
  • An append-only record of transitions and their causes makes the state auditable.Supporting

The reasoning

  1. Model payment as attempts with their own keys and an explicit 'unknown' state, not a boolean.
  2. Keep current state for decisions and an append-only ledger for history, written in one transaction.
  3. Derive the order's status from its attempts.

Two tables, two jobs:

  • payment_attempts holds current state, optimized for decisions: is this attempt still processing? What key did we use? Transitions are conditional updates; see State machines for business state.
  • payment_events holds history: one row per transition, with the cause and the provider's event or request ID. It is written in the same transaction as the state change, so they can never disagree; see Append-only logs.

The order's status is a function of its attempts: paid if any attempt succeeded and was not refunded. A unique partial index (ON payment_attempts (order_id) WHERE status IN ('processing', 'succeeded')) makes "one live attempt per order" a database invariant rather than a hope.

Tradeoffs

ChoiceGainsCosts
Attempts plus an append-only ledgerPer-attempt idempotency, an explicit unknown state, and a full audit trail.More tables, plus derived order status that every query must compute or maintain.

Stage 4 of 13 · Model

Which transitions can happen?

An attempt's states are created, processing, succeeded, failed and refunded. Information about it arrives from three directions: the provider's synchronous response, webhooks, and the reconciler. Decide which of these statements about transitions are true.

What you need to know first

A state machine lists which moves are allowed. Two rules make one robust when information arrives from several directions:

  1. Move only on facts. A decline is a fact; a timeout isn't.
  2. Move forward only. Each transition names the state it starts from, so a late or repeated message can't drag the state backwards.

In SQL, each move is UPDATE … SET status = 'succeeded' WHERE id = $1 AND status = 'processing'. See State machines for business state.

An attempt is 'succeeded'. A delayed 'processing' webhook for it arrives. With forward-only transitions, what happens?

Nothing: the transition to processing doesn't start from succeeded, so zero rows change.

The update's WHERE clause requires the earlier state. A stale event matches nothing and is harmlessly ignored.

Money coming back after a success isn't "succeeded became failed". It's a new move, succeeded → refunded (or a dispute), recorded as its own event. History is appended to, never rewritten.

The API response, a webhook and the reconciler all learn that attempt 77 succeeded, at about the same moment. Each calls the same conditional transition. What happens?

Exactly one of them changes the row from processing to succeeded; the other two match zero rows. The state is correct, the ledger has one success entry, and the side effects attached to that transition run once.

What the stage asks

Which statements about the attempt's lifecycle hold?

  1. Holds

    An attempt marked failed because the provider call timed out can later turn out to have succeeded.

    That is exactly why a timeout must not lead to failed. Only a definitive answer (a decline, an error the provider documents as final) may. A timeout leaves the attempt in processing.

  2. Fails

    A succeeded attempt may move to failed if a later webhook reports a failure.

    Succeeded is terminal for the charge. A later failed event for the same payment is an older event arriving late. Money coming back is a new transition (refunded, or a dispute), recorded as such, never an overwrite of history.

  3. Fails

    Applying each webhook's status to the attempt in the order webhooks arrive keeps the state correct.

    Arrival order is not event order. Either restrict transitions to forward moves (processing → succeeded matches only from processing), or treat each webhook as a nudge and fetch the payment's current state from the provider.

  4. Holds

    Refunded can only follow succeeded.

    You cannot return money you never captured. Enforce it in the transition's WHERE clause, not only in UI logic.

  5. Holds

    A card-declined response is a definitive answer that may move the attempt straight to failed.

    The provider has told you the outcome: nothing was charged. The buyer can start a new attempt, with a new key, using another card.

The reasoning

  1. Move only on facts: a timeout leaves an attempt processing; only a definitive answer fails it.
  2. Forward-only conditional transitions make late and duplicate messages harmless.
  3. Every source of truth uses the same transitions, so they converge on one answer.

The rule underneath all five: transitions move forward only, and only on facts.

created ──▶ processing ──▶ succeeded ──▶ refunded
                │
                └────────▶ failed      (only on a definitive answer)

Each arrow is an UPDATE … WHERE status = <from>. Every source of information (API response, webhook, reconciler) uses the same transitions. Whichever learns the truth first moves the state; the rest match zero rows and change nothing. That is how three independent, unordered, possibly duplicated channels converge on one answer; see State machines for business state.

Stage 5 of 13 · Break it

A buyer was charged twice

This is the prototype's checkout handler. It looks reasonable. Find every line that contributes to double charges or inconsistent state.

What you need to know first

Check-then-act reads a value, decides, then acts: "if the order isn't paid, charge it". Two requests running at once can both read "not paid" before either acts, and both charge.

Only an operation the database performs atomically, like inserting under a unique constraint, can decide "first or not" correctly when requests overlap.

A buyer double-clicks Pay. Both requests read status = 'pending', then each calls the provider without an idempotency key. What does the provider see?

Two unrelated charge requests, so it charges twice.

Without a shared key, the provider has no way to know the second request is a repeat of the first.

Look for these in any payment handler:

  • Is something durable written before the irreversible call? If not, a crash at that moment leaves no trace that a charge may exist.
  • Is the same key sent on every retry? If not, retries become new charges.
  • Are follow-up actions recorded with the state change? If not, a crash after the commit loses them.

What the stage asks

Select the lines responsible.

CodetypescriptPOST /orders/:id/pay
  1. 1app.post("/orders/:id/pay", async (req, res) => {
  2. 2 const order = await db.orders.find(req.params.id);
  3. 3 if (order.status === "paid") return res.json({ status: "paid" });

    Check-then-act: two concurrent requests both read 'unpaid' before either finishes. Deduplication must be an atomic claim, such as a unique idempotency key, not a read.

  4. 4 const charge = await provider.payments.create({

    Nothing is recorded before calling the provider. If this request times out or the process dies here, there is no trace that a charge may exist.

  5. 5 amount: order.total,
  6. 6 currency: order.currency,
  7. 7 paymentToken: req.body.token,

    No idempotency key is sent. The provider cannot recognize the second click (or a retry after a timeout) as the same intent, so it creates a second charge.

  8. 8 });
  9. 9 if (charge.status === "succeeded") {
  10. 10 await db.orders.update(order.id, { status: "paid" });
  11. 11 await enrollments.grant(order.userId, order.courseId);

    Enrollment and email run after the commit with nothing recording that they are owed. A crash here leaves a paid order with no course, and no process will ever finish it.

  12. 12 await email.sendReceipt(order);
  13. 13 }
  14. 14 res.json({ status: charge.status });
  15. 15});

What the fix has to do

  • The first request's outcome was unknown to the buyer; retrying without a stable idempotency key created a second charge.
  • The status check races with concurrent requests; dedupe must be an atomic claim (unique key / insert-if-absent).
  • The attempt must be durably recorded, with its key, before the provider is called.
  • Follow-up effects need to be recorded in the same transaction as the status change (outbox) so a crash cannot drop them.Supporting

The reasoning

  1. Check-then-act races: dedupe with an atomic claim such as a unique idempotency key.
  2. Record the attempt durably, with its key, before calling the provider.
  3. The unit of reliability is a durable attempt, not the request.

Two clicks, both reading status = 'pending', both calling the provider with no key. The provider saw two unrelated requests and did exactly what it was asked.

Every line flagged is a version of one mistake: acting as though the request is the unit of reliability. The request can be duplicated, cut off, or abandoned at any line. The unit of reliability has to be something durable that every duplicate can find: an attempt row, created atomically under a key, before anything irreversible happens; see Idempotency.

Stage 6 of 13 · Break it

Write the idempotent checkout

Rewrite the handler. The client sends an Idempotency-Key header that it generates once per checkout attempt and reuses on every retry of that attempt. Use whatever mix of SQL and TypeScript you like. What matters is which operations are atomic, what is recorded when, and what each path returns.

What you need to know first

The atomic claim in SQL:

INSERT INTO payment_attempts (order_id, idempotency_key, status)
VALUES ($1, $2, 'processing')
ON CONFLICT (idempotency_key) DO NOTHING
RETURNING *;

If a row comes back, this request created the attempt and should call the provider. If nothing comes back, another request already did, and this one should read and return that attempt's state.

A second request with the same key arrives while the first is still waiting on the provider. What should it do?

Return the existing attempt's current state ('processing', 202)

The first request owns the call. The second just reports what's known. If the buyer keeps polling, they'll see the result once it's applied.

The same key can be reused by mistake for a different purchase: a bug that sends a cached key with a new order. Compare the stored attempt's order and amount with the request's, and reject a mismatch (422) rather than returning the old result.

The process crashes right after inserting the attempt and before calling the provider. What state is the system in, and how does it recover?

An attempt in processing that the provider has never heard of. The buyer's retry (same key) finds it. A reconciler looking at old processing attempts asks the provider, learns there's no such payment, and once no request could still be in flight, marks it failed so the buyer can try again.

What the stage asks

Implement the checkout handler so retries, double-clicks and timeouts cannot produce a second charge.

Reference implementation

app.post("/orders/:id/pay", async (req, res) => {
  const key = req.header("Idempotency-Key");
  if (!key) return res.status(400).json({ error: "Idempotency-Key required" });
  const order = await db.orders.find(req.params.id);

  // 1. Atomic claim: exactly one request creates the attempt for this key.
  const inserted = await db.oneOrNone(`
    INSERT INTO payment_attempts (order_id, idempotency_key, amount, currency, status)
    VALUES ($1, $2, $3, $4, 'processing')
    ON CONFLICT (idempotency_key) DO NOTHING
    RETURNING *`, [order.id, key, order.total, order.currency]);

  if (!inserted) {
    const existing = await db.one(
      `SELECT * FROM payment_attempts WHERE idempotency_key = $1`, [key]);
    if (existing.order_id !== order.id || existing.amount !== order.total) {
      return res.status(422).json({ error: "Key reused with different parameters" });
    }
    return res.status(existing.status === "processing" ? 202 : 200).json(view(existing));
  }

  // 2. Call the provider with the same key. Bounded wait.
  let result;
  try {
    result = await provider.payments.create(
      { amount: order.total, currency: order.currency,
        paymentToken: req.body.token, metadata: { attempt: inserted.id } },
      { idempotencyKey: inserted.idempotency_key, timeoutMs: 8000 });
  } catch (err) {
    if (isDefinitiveDecline(err)) {
      await applyTransition(inserted.id, "failed", { cause: "api", reason: err.code });
      return res.status(200).json({ status: "failed", reason: err.code });
    }
    // Timeout or network error: the outcome is unknown. Leave it processing.
    return res.status(202).json({ status: "processing" });
  }

  await applyTransition(inserted.id, result.status, { cause: "api", providerRef: result.id });
  return res.status(result.status === "processing" ? 202 : 200).json({ status: result.status });
});

// Shared by the API, the webhook receiver and the reconciler.
async function applyTransition(attemptId, to, cause) {
  const from = { succeeded: ["processing"], failed: ["processing"], refunded: ["succeeded"] }[to];
  if (!from) return false; // "processing" again: nothing to do
  return db.tx(async (tx) => {
    const moved = await tx.result(`
      UPDATE payment_attempts SET status = $2, provider_payment_id = COALESCE($3, provider_payment_id)
      WHERE id = $1 AND status = ANY($4)`, [attemptId, to, cause.providerRef, from]);
    if (moved.rowCount === 0) return false; // already moved by another path
    await tx.none(`INSERT INTO payment_events (attempt_id, kind, cause, provider_ref)
                   VALUES ($1, $2, $3, $4)`, [attemptId, to, cause.cause, cause.providerRef]);
    if (to === "succeeded") await enqueue(tx, ["grant-enrollment", "send-receipt"], attemptId);
    if (to === "refunded") await enqueue(tx, ["revoke-enrollment"], attemptId);
    return true;
  });
}
  • The unique key and ON CONFLICT DO NOTHING make the claim atomic. Two simultaneous clicks race on an index, not on application logic, and exactly one wins.
  • The attempt is created as processing before the provider call. From that moment, a durable record says "a charge may exist". Webhooks, the reconciler and the buyer's retries all find it.
  • Timeouts never produce failed. Only a definitive answer does. Everything else is resolved later through applyTransition, the single function every source of truth uses, so they cannot conflict.
  • Side effects (grant-enrollment, send-receipt) are outbox rows inserted in the same transaction as the transition. The next stages explain why.

What a strong answer covers

  • The attempt is inserted under a unique idempotency key in one atomic statement; a duplicate finds the existing attempt instead of creating one.
  • The provider call carries the attempt's key, so the provider deduplicates retries too.
  • A timeout or network error leaves the attempt in processing and returns 202, never failed.
  • Results are applied with a conditional transition from processing, together with a ledger event and outbox rows in one transaction.
  • Reusing a key with a different order or amount is rejected.Supporting
  • A duplicate arriving while the first is still processing returns the current state rather than calling again.Supporting

The reasoning

  1. Insert the attempt under a unique key with ON CONFLICT DO NOTHING: exactly one request wins.
  2. Send the attempt's key to the provider; a timeout leaves the attempt processing, never failed.
  3. Apply every outcome through one conditional transition that also writes the ledger and outbox.

Notice what the handler no longer depends on: the request completing. If the process dies after the insert, the attempt sits in processing and gets resolved. If it dies after the provider call, the provider holds the key and will return the same result to any retry, and the webhook will arrive anyway. If the buyer clicks five times, four of the clicks read an existing row.

The request is now just one of several ways to learn an outcome that is defined by durable state.

Stage 7 of 13 · Decide

Who owns the idempotency key?

The key decides which requests count as "the same". If it is too broad, legitimate new attempts get the old answer. If it is too narrow, retries become new charges. A buyer whose card was declined must be able to pay with a different card. A buyer whose connection dropped must not pay twice.

What you need to know first

An idempotency key names an intent. Requests that are the same intent must share the key; a new intent must get a new one. Getting the scope wrong fails in one of two directions:

  • Too broad: different intents share a key, so a new attempt gets an old attempt's cached answer.
  • Too narrow: retries of one intent get different keys, so each retry is treated as new.

The key is the order id. The buyer's card is declined and they enter a different card. What happens?

The provider returns the cached decline: the new card is never tried.

Same key, so the provider treats it as a repeat of the declined attempt for 24 hours. Too broad.

The key is a hash of the request body, including the single-use payment token. The request times out, and the buyer re-enters the same card, which produces a new token. What happens?

New token, new hash, new key: if the first charge succeeded, the buyer is charged twice.

The retry is the same intent but looks different in bytes. Too narrow, which is the dangerous direction with money.

What the stage asks

What should the idempotency key identify, and who generates it?

  1. Sound

    The browser generates a UUID when the buyer starts a checkout attempt, reuses it on retries, and generates a new one after a definitive decline

    The key matches the intent: this attempt at paying. Network retries and double-clicks reuse it; a new card after a decline is a new intent and gets a new key. The server still enforces uniqueness, so a misbehaving client cannot cause a double charge, only confusing errors.

  2. Flawed

    Use the order id as the key

    Too broad. After a decline the buyer tries another card, sends the same key, and the provider, honouring its contract, returns the cached decline. The buyer can never pay for this order through this path within 24 hours.

  3. Sound

    Server-derived: order id plus an attempt counter that increments only after a definitive failure

    This works without trusting the client: while an attempt is processing, every request maps to it; once it has definitively failed, the next request opens attempt n+1. The server must make that "current attempt" decision atomically, which is what the unique partial index on live attempts gives you.

  4. Flawed

    A hash of the request body: order, amount and payment token

    Too narrow. Payment tokens are single-use, so a buyer who re-enters the same card after a timeout gets a new token, a new hash, and a second charge. A key must identify intent, not bytes.

What a strong answer covers

  • The key identifies one payment attempt: retries of it share the key, and a genuinely new attempt gets a new one.
  • Explains a concrete failure of a key that is too broad (stuck on a cached decline) or too narrow (new key on retry, so a double charge).
  • The server's own uniqueness checks must not depend solely on the client behaving.Supporting
  • Retries arriving after the provider's 24-hour retention would bypass its dedupe, so your own record must catch them.Supporting

The reasoning

  1. An idempotency key names one intent: retries share it, a genuinely new attempt gets a new one.
  2. Too broad traps buyers on cached declines; too narrow double-charges.
  3. Keep your own record of keys, since the provider forgets them after 24 hours.

An idempotency key is a name for an intent, and getting its scope right is the whole problem. There are two ways to get it wrong:

  • Too broad turns new intents into replays. The buyer is stuck on a cached decline.
  • Too narrow turns replays into new intents. The buyer is charged twice.

With money, too narrow is the dangerous direction; too broad is a support ticket.

Note the 24-hour retention. A buyer whose browser retries a request from yesterday's tab will pass the provider's dedupe. Your own attempt row, keyed by the same value and retained indefinitely, is what catches it.

Where another engineer could land differently

Native mobile apps often prefer server-derived keys, because a reinstall or a crash can lose a client-generated key that was never persisted.

Stage 8 of 13 · Break it

Webhooks arrive twice, and out of order

The provider's documentation says it all plainly: events can be delivered more than once, in any order, and your endpoint must respond within 10 seconds or the delivery is retried.

What you need to know first

A webhook is a hint that something may have changed, not an instruction. Handlers that treat it as an instruction ("set status to X, then send the receipt") break in two ways:

  • A duplicate delivery repeats the instruction: two receipts.
  • A late delivery applies an old status over a newer one.

Two ways to make a webhook handler safe:

  • Deduplicate and move forward only: record each event id under a unique constraint, skip ones already seen, and apply only forward transitions.
  • Fetch current state: treat the webhook as a nudge and ask the provider for the payment's latest state. Ordering stops mattering, at the cost of one API call per event.

Either way, verify the signature first, so nobody can forge "payment succeeded".

Where should the 'send receipt' action be triggered?

By the transition to succeeded, as an outbox row written in the same transaction

The transition happens once, however many times the event arrives. Attaching effects to it means a duplicate event, matching zero rows, triggers nothing.

A duplicate payment.succeeded event arrives for an attempt that's already succeeded. What should the handler respond, and why?

200 OK. The event was handled: there was simply nothing left to do. Responding with an error would make the provider retry it for days, for no benefit.

What the stage asks

How should the webhook receiver handle events?

  1. Flawed

    Set the attempt's status to whatever the event says, then run the side effects for that status

    This is the handler that produced the incident: the duplicate re-ran side effects (two receipts) and the late processing event moved a succeeded payment backwards.

  2. Sound

    Verify the signature, record the event id under a unique constraint, apply the transition only if it moves forward from the current state, and respond 200

    Duplicates hit the unique event-ID constraint, or match zero rows in the forward-only transition. Late events cannot regress state. Side effects are not triggered by the event at all: they are outbox rows created by the transition, so they happen exactly when the state actually changes, once.

  3. Sound

    Treat the event as a nudge: verify it, then fetch the payment's current state from the provider and apply that

    Ordering stops mattering, because you always apply the provider's latest view. The cost is one API call per webhook, trivial at 50 orders a minute but significant against a 100 requests/second limit at much larger scale. You still need forward-only transitions, because two fetches can race.

  4. Flawed

    Grant the enrollment and send the receipt inside the webhook request, then respond 200

    Slow side effects risk the 10-second deadline, a missed deadline causes a redelivery, and the redelivery runs the side effects again. A failing email provider would also make the payment provider retry webhooks for three days.

What a strong answer covers

  • Repeated events are deduplicated by event id, or the transition itself is idempotent.
  • Out-of-order events cannot regress state, either through forward-only transitions or by fetching current state.
  • Side effects are triggered by the state transition, not by receiving the event, and run outside the webhook request.
  • Signatures are verified so forged events cannot mark payments succeeded.Supporting

The reasoning

  1. Treat webhooks as hints: verify, deduplicate by event id, and apply forward-only transitions.
  2. Attach side effects to transitions (which happen once), not to messages (which repeat).
  3. Respond quickly, and with 200 for duplicates, so the provider stops retrying.

A webhook is a hint that something may have changed, not a command to do something. Handling it means:

  1. Verify the signature and timestamp.
  2. INSERT INTO processed_events (event_id) … ON CONFLICT DO NOTHING; if it already exists, respond 200 and stop.
  3. Call applyTransition, the same function the API uses, in the same transaction.
  4. Respond 200 within milliseconds.

The receipt duplication was not really a webhook bug. It was a side effect attached to the wrong thing. Attach effects to transitions, which happen once, rather than to messages, which can arrive any number of times; see Webhooks and Delivery guarantees.

Tradeoffs

ChoiceGainsCosts
Apply the event's payloadNo extra API calls.Must reason carefully about which transitions each event type may trigger.
Fetch current state on each eventImmune to ordering; simple handler logic.An API call per event, which counts against rate limits at scale.

Stage 9 of 13 · Break it

The commit nobody acted on

The payment state is right. What is missing is everything that was supposed to follow from it. The state change and the follow-up actions live in different systems.

What you need to know first

After a payment succeeds, other things must follow: grant the course, send a receipt. These live in other systems. If your code commits the payment and then calls them, a crash between the two loses the follow-up, and nothing records that it was owed.

Logging an error doesn't help: a killed process writes no log line.

The transactional outbox records the obligation inside the same transaction as the state change: "grant course for attempt 77" and "send receipt for attempt 77" as rows in an outbox table. Either the payment and its obligations commit together, or neither does.

A worker then reads outbox rows, performs each action, and marks it done. If it crashes, the row is still there and is retried. See Transactional outbox.

The outbox worker grants the course, then crashes before marking the row done. What happens next, and what must the handler do about it?

The row is retried and the course granted again, so granting must be idempotent (an upsert on user and course).

Outbox delivery is at least once. Each handler has to tolerate running twice, for example with ON CONFLICT DO NOTHING.

The worker grants the course and sends the receipt. If it can only finish one before crashing, which order leaves the buyer better off?

Grant first, then send the receipt. The worst partial state is a buyer with access and no email, which nobody notices. The reverse leaves a buyer with a receipt and no course, which is a support ticket.

What the stage asks

How should side effects follow a successful payment?

  1. Flawed

    Run them right after the commit; if one fails, log an error for an engineer

    A deploy, an OOM kill or a crash between the commit and the call leaves no log line, because the process is gone. Nothing records that work is owed, so nothing ever does it.

  2. Sound

    In the same transaction as the transition, insert outbox rows; a worker processes them at least once with idempotent handlers

    The transition and the obligation to act on it commit together, or not at all. A worker claims outbox rows, performs each effect, and marks it done; a crash means the row is retried. Handlers are idempotent: enrollment is an upsert on (user_id, course_id), and the receipt carries an idempotency key derived from the attempt.

  3. Flawed

    Publish a payment-succeeded message to a queue after the commit

    The same gap, moved: a crash between the commit and the publish loses the message. A queue is a fine transport for outbox rows, but something durable has to remember the message was owed until it is sent.

  4. Flawed

    Use a distributed transaction across Postgres, the enrollment service and the email provider

    The email provider will not join your transaction, and an email cannot be rolled back once sent. Two-phase commit needs every participant's cooperation, and it would block your database whenever any of them is slow.

What a strong answer covers

  • The state change and the notification are writes to different systems; a crash between them loses one.
  • Recording the obligation in the same transaction as the state change makes them atomic.
  • Delivery from the outbox is at-least-once, so each effect must be idempotent.
  • External effects like email cannot be rolled back, which is why they come after the commit, not inside it.Supporting

The reasoning

  1. Record follow-up actions as outbox rows in the same transaction as the state change.
  2. Outbox delivery is at least once, so every handler must be idempotent.
  3. Order effects so the least harmful partial state comes first.

The Transactional outbox turns "do these things after the commit" into "record that these things are owed, as part of the commit". The worker then makes the record true, retrying until it is.

The guarantee is eventually, at least once. That is why each handler has its own deduplication:

  • Enrollment: INSERT … ON CONFLICT (user_id, course_id) DO NOTHING.
  • Receipt: the email provider's idempotency key, set to receipt:{attempt_id}.

Grant the course before sending the receipt in the worker's sequence, so the worst case of a crash is a buyer with access but no email. Ordering effects so that partial completion is the less harmful partial state is the same habit as in every other investigation.

Stage 10 of 13 · Break it

Thirty-seven payments stuck in processing

Every component behaved as designed: the provider retried as documented, your handler rejected what it could not verify, and the attempts stayed in processing instead of guessing. Now the system needs a way to finish what messages could not.

What you need to know first

Messages can't guarantee resolution: a webhook endpoint broken for longer than the provider's retry window loses those events permanently. Reconciliation asks the source of truth directly and fixes what the messages missed. See Reconciliation.

A reconciler has no special powers. It learns a fact, then submits it through the same transitions as the API and webhooks, so it can't conflict with them.

The reconciler asks the provider about an attempt that has been processing for 20 minutes. The provider has no record of it. What should happen?

Mark it failed, once it's well past any request timeout so no request could still land.

After minutes, a request that never arrived can't arrive any more. Failing it lets the buyer try again. A retry carrying the same key would be deduplicated anyway.

Two kinds of reconciliation catch different things:

  • Targeted, every few minutes: old processing attempts, looked up one by one.
  • Full, daily: the provider's settlement report compared with your ledger in both directions: charged but not recorded, and recorded but not charged.

One metric would have caught the broken endpoint on day one: the age of the oldest processing attempt.

What the stage asks

Design the reconciliation process: what it looks for, how it decides each case, how it applies what it learns, and what it must never decide on its own.

Reference answer

Targeted sweep, every few minutes:

SELECT * FROM payment_attempts
WHERE status = 'processing' AND created_at < now() - interval '15 minutes';

For each attempt, look it up at the provider by the idempotency key or the metadata.attempt reference:

  • Succeeded or failed at the provider: call applyTransition with cause: 'reconciliation'. If a webhook gets there first, one of them matches zero rows. No conflict is possible.
  • Still processing at the provider (for example, waiting on 3-D Secure): leave it; after a policy timeout, cancel it at the provider, then fail it here.
  • No record at the provider: the request may never have arrived, or may still be in flight. Once well past your own request timeout (minutes), it cannot land, so mark it failed with the reason "not received by provider". Because the provider remembers keys for 24 hours, a retry with the same key would still deduplicate even if you were wrong.

Daily full reconciliation: compare the provider's settlement report with payment_events, in both directions. "Captured at provider, not succeeded here" and "succeeded here, not captured there" each go to a human queue with full context. These are the cases where something outside the design happened, such as manual dashboard actions or provider bugs.

Alerting: the age of the oldest processing attempt is a single metric that would have paged someone on day one instead of day four. The reconciler's output is a report on the health of every other path; see Reconciliation.

What a strong answer covers

  • Periodically finds attempts in processing older than a threshold and looks each up at the provider by its key or reference.
  • Applies the provider's authoritative state through the same conditional transitions as the API and webhooks, so they cannot conflict.
  • Treats 'provider has no record' carefully: waits until no in-flight request could still land before failing it, and escalates true anomalies to a human.
  • Also compares the daily settlement report with the ledger in both directions (charged but not recorded, recorded but not charged).Supporting
  • Alerts on the number and age of unresolved attempts, which would have caught the broken endpoint within hours.Supporting

The reasoning

  1. Reconciliation resolves what messages couldn't, by asking the source of truth.
  2. Apply what it learns through the same conditional transitions as every other path.
  3. Alert on the age of the oldest unresolved attempt, and diff settlement reports daily.

Reconciliation is not an admission that the design failed. It is the part of the design that handles the messages no design can guarantee. Webhooks make resolution fast; reconciliation makes it certain.

The essential property is that the reconciler has no special powers: it learns a fact from the source of truth and submits it through the same door as everyone else.

Stage 11 of 13 · Change it

Adding refunds

A refund is money moving the other way, through the same provider, with the same latency and the same uncertainty.

What you need to know first

A refund is money moving the other way through the same provider, with the same latency, timeouts and uncertainty as a charge. Everything that made charges safe applies again: a durable record before the call, an idempotency key, a state for "unknown", and resolution from the provider.

The handler sets the attempt to 'refunded', revokes access, then calls the provider's refund endpoint, which times out. What's wrong?

Your records say the buyer was refunded when they may not have been. The state moved on a hope rather than a fact. If the refund never happened, the buyer lost access and kept paying.

An admin double-clicks Refund. Without an idempotency key on the refund call, what can happen?

Two refunds are issued for one payment.

Each click is a separate refund request to the provider. A refund record with its own key makes the second click a repeat.

What the stage asks

How should a refund be performed?

  1. Flawed

    Set the attempt to refunded, revoke access, then call the provider's refund endpoint

    If the provider call fails or times out, your records say the buyer was refunded when they were not. The state machine has moved on a hope rather than a fact.

  2. Defensible

    Call the provider's refund endpoint; if it returns success, set the attempt to refunded

    The order is right, but there is no idempotency key, so an admin double-clicking or a retry after a timeout can issue two refunds, and nothing durable records that a refund may be in flight. It works most of the time, which is the most dangerous kind of payment code.

  3. Sound

    Record a refund request (its own row, key and state machine) in a transaction, call the provider with that key, and apply the outcome through the usual transitions

    A refund is a new money movement, so it gets the full treatment: a durable record before the call, an idempotency key, an unknown state, resolution by response, webhook or reconciliation, and access revocation as an outbox effect of the succeeded → refunded transition. Partial refunds later are just multiple refund rows.

  4. Defensible

    Have support refund in the provider's dashboard; the refund webhook updates your state

    It works at low volume, and your webhook handling and reconciliation will catch it. But nothing in your system records who refunded or why, and access revocation now depends entirely on webhook delivery. It is a fine stopgap, but not a design.

What a strong answer covers

  • A refund has the same uncertainty as a charge: it needs its own record, idempotency key and unknown state.
  • State changes to refunded only on a confirmed provider outcome.
  • History is appended (a refund event), never rewritten.Supporting

The reasoning

  1. A refund has the same uncertainty as a charge: give it its own record, key and unknown state.
  2. Move to refunded only on a confirmed provider outcome; revoke access as an outbox effect.
  3. A design that absorbs new features by reusing its own pattern captured the problem.

The design absorbed a new requirement by reusing its own pattern: new entity, key, conditional transitions, outbox effects, reconciliation. That is the test of whether a design captured the problem or just the first feature. Refunds did not need a new idea; they needed the old idea applied again.

The ledger now tells the full story: processing → succeeded (evt_A) → refund requested (admin 14) → refunded (evt_B).

Stage 12 of 13 · Change it

100x volume and a second provider

The patterns hold. The question is what new pressure appears, and which tempting shortcuts would break the guarantees.

What you need to know first

Flash sales reach 5,000 orders a minute. About how many create-payment requests a second is that?

About 83 per second.

5,000 ÷ 60 ≈ 83 a second, before any retries, lookups or refunds. The provider's limit is 100, so the limit now shapes the design: shared outbound rate limiting, and queueing for anything not interactive.

During a provider brownout, retries arrive exactly when the provider can least absorb them, and they use up your own rate limit. Bound them: exponential backoff with jitter, a retry budget, and a circuit breaker that stops new charges and tells buyers clearly. See Retries, backoff and jitter.

Provider A times out on a charge. Is it safe to immediately retry the same charge with provider B?

No: A's outcome is unknown, and the buyer may already have been charged.

A new attempt with B is only safe once A's attempt is definitively failed. Otherwise failover creates the double charge the design exists to prevent.

What the stage asks

Which statements hold at the new scale?

  1. Holds

    During flash sales, the provider's 100 requests/second limit becomes a real constraint.

    5,000 orders a minute is about 83 creates per second before any retries, lookups or refunds. You need a shared outbound limiter, queueing for non-interactive calls, and a conversation with the provider about raising the limit. See Rate limiting.

  2. Holds

    During a provider brownout, aggressive retries can make the outage worse for everyone.

    Retries multiply load exactly when the provider is least able to absorb it, and they consume your own rate limit. Use backoff with jitter, a retry budget, and a circuit breaker that stops sending new charges and shows buyers a clear message. See Retries, backoff and jitter.

  3. Fails

    If provider A times out, immediately retrying the charge with provider B is a safe failover.

    A's outcome is unknown: the buyer may already have been charged. Failing over without resolving it creates exactly the double charge the design exists to prevent. Each attempt is bound to one provider; a new attempt with B is only allowed once A's attempt is definitively failed or cancelled.

  4. Fails

    Supporting two providers requires a distributed transaction between them.

    Each attempt goes to exactly one provider, recorded on the attempt row along with its key. There is nothing to coordinate between providers, only per-attempt state, as before.

  5. Fails

    Postgres becomes the first bottleneck at this volume.

    Peak is roughly 100-300 small writes per second across attempts, events and outbox rows, comfortably within one Postgres primary. The provider's rate limit binds long before your database does.

The reasoning

  1. At scale the provider's rate limit binds first; add a shared outbound limiter and circuit breakers.
  2. Never fail over or retry elsewhere while an outcome is unknown.
  3. Bind each attempt to one provider and reconcile each provider separately.

At scale, the biggest risk is not throughput. It is operational shortcuts that bypass uncertainty: failing over on a timeout, retrying harder during an outage, or treating "slow" as "failed" to keep queues moving. Each is tempting during an incident, and each reintroduces double charges.

The scalable additions are the boring ones: a shared outbound rate limiter, circuit breakers, per-provider routing recorded on the attempt, and reconciliation run separately for each provider.

Stage 13 of 13 · Defend it

Defend the guarantee

A new engineer joins the payments team and asks: "How do we know we can't double-charge? And is there any way we still could?"

Answer in one page, the way you would in a design review or an interview.

What you need to know first

A convincing guarantee walks through each way the bad outcome could happen and names the mechanism that stops it:

ThreatMechanism
Double-clickunique key on the attempt
Retry after timeoutsame key sent to the provider
Crash mid-flightattempt stays processing; resolved later
Duplicate or late webhookforward-only transitions; effects on transitions

Then it names the paths the mechanisms don't cover.

Which of these is a real residual risk to state honestly?

A charge made in the provider's dashboard or a script bypasses the attempt table; only reconciliation catches it.

The guarantees cover money that moves through the attempt table. Anything that moves around it is caught after the fact, which is worth saying out loud.

What the stage asks

Explain why the system cannot double-charge, which mechanism covers each failure, and the residual risks it does not cover.

Reference answer

The mechanism. A charge can only be created by the checkout handler, and only after it has inserted a payment_attempts row under a unique idempotency key. Every provider call for that attempt carries the same key. So:

  • Double-clicks and client retries hit the unique index: one request inserts, the others read the existing attempt.
  • Retries after a timeout reuse the attempt's key, so the provider returns the original result instead of charging again.
  • Crashes at any point leave a processing attempt, which webhooks or the reconciler resolve; nothing ever needs to "try again from scratch".
  • Duplicate or late webhooks cannot regress or repeat anything, because transitions are forward-only conditional updates and side effects hang off transitions via the outbox.

Residual risks, stated honestly:

  • A client that generates a new key for what is really a retry is indistinguishable from a new attempt; the unique partial index on live attempts per order is the second line of defence.
  • A retry arriving more than 24 hours later passes the provider's dedupe; our own attempt row catches it only if it carries the same key.
  • Charges made outside the system (dashboard, scripts) bypass all of this. Daily reconciliation detects them after the fact.
  • Failing over to a second provider while the first is unresolved would break the guarantee. It is forbidden by design, and that rule needs defending in every incident review.

What is not guaranteed: that the outcome is known immediately (it may take minutes), or that a receipt email is sent exactly once (it is sent at least once, deduplicated by the email provider where possible).

What a strong answer covers

  • Every charge is made under an idempotency key belonging to an attempt that is durably recorded before the call.
  • Duplicate requests, retries and events converge through unique constraints and conditional transitions.
  • Unknown outcomes are resolved from the source of truth (webhooks, lookup, reconciliation), never guessed.
  • Names residual risks honestly, e.g. retries after the provider's key retention, manual dashboard actions, failover to another provider, or a client generating new keys on retry.
  • Distinguishes the guarantee (no double charge) from what is not guaranteed (instant resolution, emails exactly once).Supporting

The reasoning

  1. Defend a guarantee threat by threat, naming the mechanism that covers each.
  2. Name residual risks: manual charges, late retries past key retention, failover, clients that mint new keys.
  3. Distinguish what's guaranteed (no double charge) from what isn't (instant resolution).

The strongest defence of a design names its own limits. "It can't double-charge" invites disbelief; "it can't double-charge through any path that goes through the attempt table, and here are the three paths that don't" invites trust, and tells the reader exactly where to look during an incident.

How you did

Now try it as an interview question

  • “Design a payment system for an e-commerce checkout.”
  • “How do you make an API endpoint idempotent?”
  • “Your service called a payment API and the request timed out. What do you do?”
  • “Design a wallet or ledger service that must never lose or duplicate money.”
  • “How would you handle webhooks from Stripe reliably?”

The interview mode mixes stages from this and other investigations with concept recall and questions about your own projects.

Back to the last stage