Append-only logs
Recording changes as an ordered, immutable sequence of facts, from which current state, history and replicas can be derived.
Storage & state
Learn it
A row updated in place answers "what's the state now?" but not "how did it get here?", "what was it at 3 p.m.?" or "what changed since I last looked?".
An append-only log records each change as an entry with a position:
(document 42, seq 5131, op …). Entries are never modified. Current state is a fold over the log: start empty and apply entries in order.What a log buys:
- History and audit: every change, when, and what caused it.
- Catch-up by position: a client or replica that has seen up to position n asks for everything after n. No per-client buffers. Database replication, Kafka consumers and collaborative editors all resume this way.
- Idempotent consumers: a consumer that remembers its position can safely re-read.
Check
A client was offline and last saw position 900. The log is at 1,250. What does it need from the server?What a log costs:
- Replay time grows forever. Periodic snapshots store the folded state as of an exact position; loading is "latest snapshot, plus entries after it".
- Ordering needs a single sequencer per log, or a partition per key; see Ordering.
- Corrections are new entries. A refund is a new event, not a deletion of the charge.
Think first
A snapshot was taken 'around position 5,000' but the exact position wasn't recorded. What goes wrong on load?
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Snapshot without its position
- Replay starts at the wrong entry, skipping or re-applying changes.
- Non-deterministic apply
- Entries whose effect depends on wall-clock time or external calls replay differently.
- Two writers, one log
- Two processes assign the same position; history forks. A unique constraint on (log, position) catches it.
Instead, consider
- Mutable rows plus an audit table
- You need history for compliance but not catch-up or replay. Simpler, and common.
- Periodic full snapshots only
- Coarse history is enough and changes between snapshots do not matter.
In practice
- Table with (stream_id, seq) primary key
- The simplest durable log, in the database you already have.
- Database write-ahead logs
- The same idea used internally for crash recovery and replication.
- Kafka, Kinesis, EventStoreDB
- Dedicated log stores with retention and consumers.
It assumes
- There is a single authority assigning positions within each log (or partition).
- Entries are deterministic to apply, so replaying the same log yields the same state.
- Storage growth is managed with snapshots, compaction or retention.
Explain it in your own words
Where you practise it
- A URL shortener like bit.ly
- A job queue that keeps working when workers fall behind
- Product analytics over billions of events
The database is down for twenty minutes · Write the batch inserter
- A payment workflow that never double-charges
- A look-aside cache at Facebook's scale
- A real-time collaborative editor
Acknowledged, then lost · Write the reconnect protocol · History and fast loads
- Sharding Postgres while it is running
Further reading
Engineers describing it in systems they run.
- How Convex Works
Convex · Sujay Jayakar · Post, Apr 2024
A walk through a reactive database from the inside: the transaction log, read sets, optimistic concurrency, and how a write finds the subscriptions it affects.
- Upgrading short link analytics by 100x with Steven Tey
Dub · Steven Tey · Post, May 2023
An interview on Tinybird's blog about why click data outgrew Redis sorted sets once people wanted to filter and group it, and what replaced them.
- Making multiplayer more reliable
Figma · Darren Tsung · Post, Oct 2022
Closing the gap between periodic saves: a durable journal of small changes, and one owner per document's history.
- Scaling Slack's Job Queue
Slack · Saroj Yadav and others · Post, Dec 2017
An outage, the design flaw behind it, and the redesign. The section on rolling the new path out without an outage is worth reading twice.
Related concepts
- Ordering
There is no global 'now' in a distributed system. Order exists only where something assigns it, so decide which order you need and who assigns it.
- Durability
What has to have happened before a system may say "saved": which failures the data must survive, and where that guarantee is actually made.
- Conflict resolution and convergence
When replicas accept concurrent changes, a deterministic rule must merge them so every replica ends in the same state without losing intent.
- Reconciliation
Periodically comparing your records with an authoritative source and repairing differences, the backstop for every message that was lost.
- Online data migrations
Moving live data to a new schema or store without downtime: write to both, backfill the past, verify, switch reads, then switch writes, with a way back at every step.
- Log-structured storage (LSM trees)
Storage engines that turn every write into a sequential append and merge files in the background: very fast writes, at the cost of compaction, tombstones and more expensive reads.
- Columnar storage
Storing each column of a table separately, so analytical queries read only the columns they use and compress them well, at the cost of slow single-row lookups and updates.