Generating unique identifiers
Making ids that are unique across machines and time, and choosing what else they reveal: order, volume, guessability, length.
Distribution
Learn it
Every record needs an id that no other record has. In one database, an auto-increment column does it. When many servers create records at once, you need a scheme, and uniqueness isn't the only thing it decides: ids also reveal information, take up space in URLs, and affect how fast inserts are.
Scheme How it stays unique What it reveals Auto-increment One counter hands out numbers How many you've made, and every other id Random (UUIDv4, random base62) The space is so large that repeats are rare Nothing Time-ordered (Snowflake, UUIDv7, ULID) Timestamp plus machine id or randomness When it was created Hash of the content The same content always gets the same id That two records have the same content Check
Invoice URLs look like/invoices/10482. A customer sees invoice 10482 one Monday and 10982 the next. What can they work out?Work it out
You've issued 1 billion random 7-character base62 codes, out of about 3.5 trillion possible. The chance that the next random code is already taken is about 1 in how many?So collisions will happen. What matters is what happens when one does. Back every scheme with a unique constraint: a collision becomes a failed insert followed by a retry with a new id, never an overwritten record.
Check
Why can random UUIDv4 primary keys slow down inserts on a very large table, compared with time-ordered ids?
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Enumeration
- Sequential public ids let anyone walk through every resource or estimate your volume.
- Silent overwrite on collision
- An upsert keyed on a colliding id replaces someone else's record.
- Clock-based ids going backwards
- Time-ordered generators must handle clock adjustments, or they emit duplicates or out-of-order ids.
- Weak randomness
- Math.random-style generators are predictable; ids meant to be unguessable are not.
Instead, consider
- Natural keys
- The entity already has a stable, unique identity (an email for accounts, an ISBN) that will never need to change.
- Composite keys
- Uniqueness is only needed within a parent, such as (document_id, seq).
In practice
- Database identity columns
- Simplest; centrally ordered.
- UUIDv4 / UUIDv7
- Random or time-ordered 128-bit ids; v7 is index-friendly.
- Snowflake-style ids
- 64-bit: timestamp, worker id, per-millisecond sequence.
- Random base62 + unique constraint
- Short, unguessable public codes with retry on collision.
It assumes
- A unique constraint backs every scheme, so even a 'can't happen' collision fails loudly rather than overwriting data.
- Random ids use a cryptographically secure generator when guessability matters.
Explain it in your own words
Where you practise it
Further reading
Engineers describing it in systems they run.
- How Discord Stores Billions of Messages
Discord · Stanislav Vishnevskiy · Post, Jan 2017
Choosing a database and a partition key for chat history, and the surprises that followed: tombstones, and an edit racing a delete.
Related concepts
- Partitioning
Splitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.
- Concurrency control
Making read-decide-write sequences safe when other actors may change the same data in between: locks, conditional writes and constraints.
- Ordering
There is no global 'now' in a distributed system. Order exists only where something assigns it, so decide which order you need and who assigns it.