Skip to content

Object storage

A flat namespace of immutable blobs addressed by key, built for durability and size rather than queries or in-place updates.

Storage & state

Learn it

0 of 2 checks done
  1. Videos, images, backups and exports are large and written once. In a relational database they bloat backups, replication and memory. On a server's local disk they're tied to a machine that will eventually be replaced.

    Object storage keeps objects (a byte blob plus metadata) under a key in a bucket. The interface is deliberately small: PUT a whole object, GET it (or a byte range), DELETE it, LIST keys by prefix. There are no in-place edits; changing an object means writing a new one.

  2. Check

    Can you append log lines to an existing object?

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Database and store disagree
A row points at a key that was never written, or objects exist that no row references. Order writes so failures leave garbage, not dangling pointers.
Abandoned multipart uploads
Uncompleted parts are invisible but billed until a lifecycle rule aborts them.
LIST as an index
Listing a large bucket to find work is slow, expensive and eventually unworkable.
Leaked presigned URLs
A URL with a long expiry grants access to anyone who obtains it.

Instead, consider

Database BLOB columns
Objects are small (kilobytes) and must change transactionally with other rows.
Block storage or a filesystem volume
Software needs POSIX semantics: random writes, appends, file locks.
Local disk
The data is scratch space that can be lost with the machine.

In practice

Amazon S3, Google Cloud Storage, Azure Blob
The reference implementations.
Cloudflare R2, Backblaze B2
S3-compatible APIs with different egress pricing.
MinIO, Ceph
Self-hosted S3-compatible stores.

It assumes

  • Objects are written whole and rarely modified; access is by key rather than by query.
  • Your database holds the metadata (which objects exist, who owns them, what state they are in).
  • Latency of tens of milliseconds per request is acceptable.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Durability

    What has to have happened before a system may say "saved": which failures the data must survive, and where that guarantee is actually made.

  • Caching

    Keeping a copy of data closer to where it is used, trading freshness and complexity for speed and reduced load on the source.

  • Asynchronous processing

    Accepting a request, recording the work durably, and doing it later in another process, so the work can outlive the request.