Skip to content

System design interview failure scenarios: 7 practice drills

Failure questions test whether your design still meets its promises after a component, request, or dependency behaves unexpectedly. Practise the reasoning in order: name the impact, protect the invariant, contain the failure, recover safely, and say what the recovery costs. These are original practice scenarios, not claims about questions asked by a particular employer.

This guide was drafted with AI assistance and has not yet been independently reviewed by a human. How SysGeeks uses AI

A five-step method for answering failure questions

  1. 1

    Describe the impact

    What does the user see? Is the request delayed, duplicated, lost, or answered with stale data?

  2. 2

    State the invariant

    What must remain true? For example, one logical payment is charged at most once, or an acknowledged write survives failover.

  3. 3

    Find the failure boundary

    Trace the request, message, or state change to the point where success becomes uncertain. Separate what is durable from what is only in memory.

  4. 4

    Contain and recover

    Explain detection, limits on retries or blast radius, the safe recovery step, and how a stale worker or duplicate event is prevented from corrupting state.

  5. 5

    Verify and name the cost

    Say which metric or reconciliation check confirms recovery, then state the latency, availability, freshness, or operational cost of your choice.

This is a reasoning aid, not a script to recite. Ask which guarantee the interviewer cares about before reaching for a queue, cache, retry, or failover policy.

Worked example: the payment response disappears

A checkout request reaches a payment provider. The provider captures the money, but the response times out before your service receives it. The customer retries. The key fact is that a timeout describes what the caller observed; it does not prove the provider did nothing.

Invariant
One logical checkout must not create two successful charges.
Containment
Keep the checkout pending and reuse its durable idempotency key for the retry.
Recovery
Query or retry against the provider using that key; reconcile cases that remain uncertain.
Trade-off
The customer may wait longer for a confirmed result, but the system avoids blindly creating another charge.

A strong answer also asks whether the provider guarantees idempotency and for how long. If its key expires before a delayed retry, the design needs a status lookup or a human reconciliation path. Practise the full payment failure.

Seven system design failure scenarios to practise

  1. The payment succeeded, but the caller timed out

    Failure: The provider may have charged the customer even though your service never received the response.

    Reason through it: Do not issue a fresh charge just because the first response is missing. Retry with the same idempotency key, query the provider by that key where supported, and reconcile outcomes that remain uncertain. Keep the order pending until the payment state is known.

    Listen for: The retry must identify the same logical payment, not create a second one.

    Practise the double-charge failure →
  2. A worker dies halfway through a job

    Failure: The job may be retried while the old worker is still running or after it has already produced an output.

    Reason through it: Make progress durable at safe boundaries, let a lease expire so another worker can take over, and attach a fencing token or attempt version to writes. Publish the result only if the worker still owns the current attempt; make repeating completed steps harmless.

    Listen for: A lease says who may work now; fencing prevents an expired owner from publishing late.

    Practise worker failure recovery →
  3. A queue delivers the same message twice

    Failure: The consumer may commit its work and crash before acknowledging the message.

    Reason through it: Assume delivery is at least once. Give each logical message a stable identifier, commit the state change and deduplication record atomically, then acknowledge. If the effect is external, use an idempotency key or an outbox-driven retry instead of claiming the whole workflow is exactly once.

    Listen for: Place the acknowledgement after durable work, and make redelivery safe.

    Practise designing a job queue →
  4. The database primary fails after acknowledging a write

    Failure: A promoted replica might not contain a write the client was told had succeeded.

    Reason through it: Ask what an acknowledgement promises and what recovery point objective is acceptable. Synchronous replication can preserve acknowledged writes at higher latency; asynchronous replication can lose recent writes during failover. Make client retries idempotent and explain how stale reads behave while replicas catch up.

    Listen for: State the durability and availability contract before choosing failover behavior.

    Review replication and failover guarantees →
  5. A cache serves stale data after an update or deletion

    Failure: A successful database write does not automatically invalidate every cached copy.

    Reason through it: Name the freshness promise first. Use versioned values or targeted invalidation for updates, bounded TTL as a backstop, and a source-of-truth check where stale data would violate a safety or deletion requirement. Explain what happens when invalidation is delayed or lost.

    Listen for: TTL bounds staleness; it does not make a cache strongly consistent.

    Practise distributed cache trade-offs →
  6. One hot key overloads a partition

    Failure: Average traffic looks healthy while one tenant, account, or object receives a disproportionate share.

    Reason through it: Measure load by key and partition, not only cluster average. Depending on the access pattern, split or salt the hot key, replicate reads, or move exceptional tenants to dedicated capacity. Then account for merge work, ordering, and the harder writes that sharding creates.

    Listen for: Scaling the cluster does not help if every request still lands on the same key.

    Practise uneven feed traffic →
  7. A live data migration misses writes during cutover

    Failure: A backfill can copy a consistent snapshot while newer writes continue to the old store.

    Reason through it: Capture changes before or during the backfill, replay them to the destination, compare counts and critical invariants, and switch reads only after the new copy catches up. Keep a rollback path and delay stopping the old write path until the new one is verified.

    Listen for: A snapshot alone cannot account for writes that arrive while it is copied.

    Practise a safe live data migration →

Check your answer before moving on

  • Did you describe what the user or caller observes, rather than only naming a component failure?
  • Did you state which guarantee must survive and identify where the operation becomes uncertain?
  • Are retries bounded and safe to repeat, including after a partial success?
  • Did you cover detection, recovery, and a way to verify the recovered state?
  • Did you name the cost or weaker guarantee your solution accepts?

For a broader practice plan, continue with the system design interview preparation guide, or use the printable interview scorecard to choose one skill to improve next.

Frequently asked questions

How should I answer a failure scenario in a system design interview?
Start with the user-visible symptom and the guarantee that must still hold. Trace the failure boundary, explain how it is detected and contained, then describe recovery and how you verify the system is correct afterward. State the cost or weaker guarantee your recovery choice introduces.
What failure scenarios should I practise for system design interviews?
Practise uncertain payment timeouts, duplicate queue delivery, workers stopping mid-job, database failover, stale caches, hot partitions, and data migrations. They test different guarantees: idempotency, durable progress, availability, freshness, load distribution, and safe cutover.
Should I retry whenever a system design request fails?
No. First decide whether the operation may already have succeeded. Retry only when the operation is safe to repeat or protected by an idempotency mechanism, and use a deadline, bounded attempts, and backoff. Otherwise check the authoritative state or reconcile the uncertain outcome.
Do interviewers expect one exact failure-handling design?
Usually the useful signal is the reasoning, not one memorized architecture. State the required guarantee, compare an appropriate recovery choice with an alternative, and explain what changes if the interviewer changes the consistency, cost, or availability requirement.

Was this useful for your system design interview prep?