Caching
Keeping a copy of data closer to where it is used, trading freshness and complexity for speed and reduced load on the source.
Performance & scale
Learn it
A cache keeps a copy of data somewhere faster or closer than the original: in the app's memory, in Redis, or on a CDN server near the user. A read checks the cache first. A hit is answered from the copy. A miss goes to the source, and usually stores the answer for next time.
Work it out
A database read takes 10 ms and a cache read takes 1 ms. A miss costs both (check the cache, then read the database). With a 90% hit rate, what's the average read time?Check
Same timings, but each item is read about once, so the hit rate is 5%. What does the cache do?Every cache is a second copy, and the copy can fall out of date. Two ways to limit that:
- TTL (time to live): copies expire after a set time, so they're never more than that old.
- Invalidation: when the data changes, delete or update the cached copy. Data is fresher, but this is easy to get wrong: a delete can be lost, or it can race with a read that puts the old value back.
Think first
Request B misses the cache and reads a user's name from the database. Before B writes it to the cache, request A renames the user in the database and deletes the cache entry. Then B writes what it read into the cache. What's in the cache now?The most reliable fix is to never change cached data at all. Give each version a new key, like
/assets/app.3f9a1c.jsoravatar-v2.png. Old copies are never wrong, only unused, so they can be cached for a year.Check
A very popular item's cache entry expires, and thousands of requests miss at the same moment and all query the database. What helps?One last rule: the system must still work with an empty cache, even if slowly. Caches get restarted and flushed. If the database can't survive the load of an empty cache, the cache has become critical infrastructure without any of a database's durability.
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Stale reads after writes
- A user updates data and immediately sees the old version.
- Thundering herd
- A hot key expires and every request hits the source simultaneously.
- Cache as source of truth
- Data that exists only in the cache is lost on eviction or restart.
- Purge does not reach every edge
- Deleted or private content remains cached at CDN edges.
Instead, consider
- Read replicas
- Staleness of seconds is fine and queries are too varied to cache by key.
- Precomputation / materialized views
- The expensive part is computing the result, and it can be refreshed on write.
- Make the source faster
- An index or a query fix removes the need for a second copy.
In practice
- In-process LRU
- Fastest; per-instance and inconsistent across instances.
- Redis / Memcached
- Shared across instances; a network hop.
- CDN
- Edge caching for static and immutable content; purge APIs for removal.
- HTTP caching headers
- Cache-Control, ETag; immutable for versioned assets.
It assumes
- Reads substantially outnumber writes for the cached data.
- Bounded staleness is acceptable, or invalidation is reliable.
- The source can survive the load of a cold cache.
Explain it in your own words
Where you practise it
- A URL shortener like bit.ly
Choose the redirect status: 301 or 302 · Make redirects fast far from the servers · Take down a malicious link everywhere · New links show as "not found" in a new region · Defend the simple design
- A reliable video processing pipeline
Trace the happy path end to end · Deleting a video mid-transcode
- A home timeline at 300,000 reads a second
What the numbers say · When does the multiplication happen? · Timelines nobody reads · The post that wouldn't die · Top posts first · Defend the hybrid
- Product analytics over billions of events
- Notifications across email, push and in-app
- A look-aside cache at Facebook's scale
How a look-aside cache behaves · A stale value that never leaves · The key everyone wants · A cache server dies · Deletes for every cluster · A second region · A cluster with an empty cache · Defend 'best-effort eventual consistency'
- Live queries: screens that update in real time
Which queries did that write change? · One board, a hundred thousand viewers · Defend building it into the database
Further reading
Engineers describing it in systems they run.
- Dub's link redirect middleware
Dub · Steven Tey and Dub contributors · Code, Sep 2022
The code that serves every Dub short link. Look for the Redis lookup with a database fallback, the click recorded after the response is sent, and what happens when Redis is failing over.
- How we built rate limiting capable of scaling to millions of domains
Cloudflare · Cloudflare · Post, Jun 2017
Approximating a sliding window with two counters, counting per data centre instead of globally, and measuring how wrong the approximation actually is.
- Timelines at Scale
Twitter (X) · Raffi Krikorian · Talk, Apr 2013
A talk on how home timelines were served: fan-out on write into a cache, what that costs for very large accounts, and why timelines are treated as rebuildable.
- Scaling Memcache at Facebook
Meta (Facebook) · Rajesh Nishtala and others · Paper, Apr 2013
The reference on running a look-aside cache hard: leases, invalidation from the commit log, failover without hammering the database, and consistency across regions.
Related concepts
- Object storage
A flat namespace of immutable blobs addressed by key, built for durability and size rather than queries or in-place updates.
- Partitioning
Splitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.
- Backpressure and capacity
When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.
- Replication
Keeping copies of data on several machines for durability, read capacity and locality, and living with copies that briefly disagree.
- Soft state
State that expires unless refreshed. It is cheap to keep, safe to lose, and right for presence, sessions and anything that describes the present moment.
- Consistent hashing
Mapping keys to nodes so that adding or removing a node moves only a small share of keys, instead of reshuffling almost all of them.
- Fan-out on write and fan-out on read
When one write must reach many readers, do the work when it is written (precompute every reader's view) or when it is read (assemble it on demand). Most real feeds do both.
- Request coalescing
When many callers ask for the same thing at the same moment, do the expensive work once and give every caller the result. The fix for thundering herds and hot keys.