Amazon: dynamo's always-on key-value store
Dynamo keeps writes available during failures with consistent hashing, replicated data, quorum-style reads and writes, and application-assisted conflict resolution.
The idea
Dynamo was designed for services that had to keep accepting reads and writes even when parts of a large data centre were unreachable. It spreads keys across a ring of nodes, replicates each item, and lets each operation choose how many replicas must respond. During a failure, a replica outside the usual preference list can temporarily accept a write so availability does not depend on every normal replica being reachable.
That choice moves complexity to recovery. Concurrent writes can produce multiple versions, so applications help reconcile them; background anti-entropy repairs replicas that drifted apart. Dynamo is a good case study in making availability a requirement with explicit consistency costs, rather than claiming both without explaining the mechanism.
The paper describes Amazon's published storage design, not a claim that Dynamo itself appears in Amazon interviews. Practise the trade-offs with the linked exercises on replication, partitioning, and recovery.
Read the originals
Written by the engineers who built it.
- Dynamo: Amazon's Highly Available Key-value Store
Giuseppe DeCandia and others · Paper, 2007
The original SOSP paper connects consistent hashing, replication, tunable read/write quorums, hinted handoff, anti-entropy, and application-assisted conflict resolution into one availability design.
Is this an actual Amazon interview question?
No. This case study explains published engineering work; it does not claim that Amazon asks this exact system in interviews. Use the source and linked exercise to practise the design decisions, constraints, and failure cases that transfer to unfamiliar interview prompts.
Practise it
Make the decisions yourself, then compare.
- A look-aside cache at Facebook's scaleDesign a Distributed Cache (Memcache)Built from Facebook's paper on scaling memcache: a look-aside cache serving billions of reads a second, Reads are fast; the hard parts are stale values, stampedes on popular keys, dead servers and invalidations that have to cross regions.Advanced · 50 min9 stages
- Sharding Postgres while it is runningShard a Live Database Without DowntimeBuilt from how Notion sharded Postgres while millions of people were using it: choose a shard key, choose a shard count you can live with for years, move every row while writes continue, prove the copy is right, and grow again later without starting over.Advanced · 50 min9 stages