Skip to content

Caching

Keeping a copy of data closer to where it is used, trading freshness and complexity for speed and reduced load on the source.

Performance & scale

Learn it

0 of 4 checks done
  1. A cache keeps a copy of data somewhere faster or closer than the original: in the app's memory, in Redis, or on a CDN server near the user. A read checks the cache first. A hit is answered from the copy. A miss goes to the source, and usually stores the answer for next time.

  2. Work it out

    A database read takes 10 ms and a cache read takes 1 ms. A miss costs both (check the cache, then read the database). With a 90% hit rate, what's the average read time?
    ms

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Stale reads after writes
A user updates data and immediately sees the old version.
Thundering herd
A hot key expires and every request hits the source simultaneously.
Cache as source of truth
Data that exists only in the cache is lost on eviction or restart.
Purge does not reach every edge
Deleted or private content remains cached at CDN edges.

Instead, consider

Read replicas
Staleness of seconds is fine and queries are too varied to cache by key.
Precomputation / materialized views
The expensive part is computing the result, and it can be refreshed on write.
Make the source faster
An index or a query fix removes the need for a second copy.

In practice

In-process LRU
Fastest; per-instance and inconsistent across instances.
Redis / Memcached
Shared across instances; a network hop.
CDN
Edge caching for static and immutable content; purge APIs for removal.
HTTP caching headers
Cache-Control, ETag; immutable for versioned assets.

It assumes

  • Reads substantially outnumber writes for the cached data.
  • Bounded staleness is acceptable, or invalidation is reliable.
  • The source can survive the load of a cold cache.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Dub's link redirect middleware

    Dub · Steven Tey and Dub contributors · Code, Sep 2022

    The code that serves every Dub short link. Look for the Redis lookup with a database fallback, the click recorded after the response is sent, and what happens when Redis is failing over.

  • How we built rate limiting capable of scaling to millions of domains

    Cloudflare · Cloudflare · Post, Jun 2017

    Approximating a sliding window with two counters, counting per data centre instead of globally, and measuring how wrong the approximation actually is.

  • Timelines at Scale

    Twitter (X) · Raffi Krikorian · Talk, Apr 2013

    A talk on how home timelines were served: fan-out on write into a cache, what that costs for very large accounts, and why timelines are treated as rebuildable.

  • Scaling Memcache at Facebook

    Meta (Facebook) · Rajesh Nishtala and others · Paper, Apr 2013

    The reference on running a look-aside cache hard: leases, invalidation from the commit log, failover without hammering the database, and consistency across regions.

  • Object storage

    A flat namespace of immutable blobs addressed by key, built for durability and size rather than queries or in-place updates.

  • Partitioning

    Splitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.

  • Backpressure and capacity

    When work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.

  • Replication

    Keeping copies of data on several machines for durability, read capacity and locality, and living with copies that briefly disagree.

  • Soft state

    State that expires unless refreshed. It is cheap to keep, safe to lose, and right for presence, sessions and anything that describes the present moment.

  • Consistent hashing

    Mapping keys to nodes so that adding or removing a node moves only a small share of keys, instead of reshuffling almost all of them.

  • Fan-out on write and fan-out on read

    When one write must reach many readers, do the work when it is written (precompute every reader's view) or when it is read (assemble it on demand). Most real feeds do both.

  • Request coalescing

    When many callers ask for the same thing at the same moment, do the expensive work once and give every caller the result. The fix for thundering herds and hot keys.