Skip to content

Append-only logs

Recording changes as an ordered, immutable sequence of facts, from which current state, history and replicas can be derived.

Storage & state

Learn it

0 of 2 checks done
  1. A row updated in place answers "what's the state now?" but not "how did it get here?", "what was it at 3 p.m.?" or "what changed since I last looked?".

    An append-only log records each change as an entry with a position: (document 42, seq 5131, op …). Entries are never modified. Current state is a fold over the log: start empty and apply entries in order.

  2. What a log buys:

    • History and audit: every change, when, and what caused it.
    • Catch-up by position: a client or replica that has seen up to position n asks for everything after n. No per-client buffers. Database replication, Kafka consumers and collaborative editors all resume this way.
    • Idempotent consumers: a consumer that remembers its position can safely re-read.
  3. Check

    A client was offline and last saw position 900. The log is at 1,250. What does it need from the server?

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Snapshot without its position
Replay starts at the wrong entry, skipping or re-applying changes.
Non-deterministic apply
Entries whose effect depends on wall-clock time or external calls replay differently.
Two writers, one log
Two processes assign the same position; history forks. A unique constraint on (log, position) catches it.

Instead, consider

Mutable rows plus an audit table
You need history for compliance but not catch-up or replay. Simpler, and common.
Periodic full snapshots only
Coarse history is enough and changes between snapshots do not matter.

In practice

Table with (stream_id, seq) primary key
The simplest durable log, in the database you already have.
Database write-ahead logs
The same idea used internally for crash recovery and replication.
Kafka, Kinesis, EventStoreDB
Dedicated log stores with retention and consumers.

It assumes

  • There is a single authority assigning positions within each log (or partition).
  • Entries are deterministic to apply, so replaying the same log yields the same state.
  • Storage growth is managed with snapshots, compaction or retention.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • How Convex Works

    Convex · Sujay Jayakar · Post, Apr 2024

    A walk through a reactive database from the inside: the transaction log, read sets, optimistic concurrency, and how a write finds the subscriptions it affects.

  • Upgrading short link analytics by 100x with Steven Tey

    Dub · Steven Tey · Post, May 2023

    An interview on Tinybird's blog about why click data outgrew Redis sorted sets once people wanted to filter and group it, and what replaced them.

  • Making multiplayer more reliable

    Figma · Darren Tsung · Post, Oct 2022

    Closing the gap between periodic saves: a durable journal of small changes, and one owner per document's history.

  • Scaling Slack's Job Queue

    Slack · Saroj Yadav and others · Post, Dec 2017

    An outage, the design flaw behind it, and the redesign. The section on rolling the new path out without an outage is worth reading twice.

  • Ordering

    There is no global 'now' in a distributed system. Order exists only where something assigns it, so decide which order you need and who assigns it.

  • Durability

    What has to have happened before a system may say "saved": which failures the data must survive, and where that guarantee is actually made.

  • Conflict resolution and convergence

    When replicas accept concurrent changes, a deterministic rule must merge them so every replica ends in the same state without losing intent.

  • Reconciliation

    Periodically comparing your records with an authoritative source and repairing differences, the backstop for every message that was lost.

  • Online data migrations

    Moving live data to a new schema or store without downtime: write to both, backfill the past, verify, switch reads, then switch writes, with a way back at every step.

  • Log-structured storage (LSM trees)

    Storage engines that turn every write into a sequential append and merge files in the background: very fast writes, at the cost of compaction, tombstones and more expensive reads.

  • Columnar storage

    Storing each column of a table separately, so analytical queries read only the columns they use and compress them well, at the cost of slow single-row lookups and updates.