Design YouTube: a system design interview walkthrough
The strongest answer separates video upload and processing from playback. This walkthrough works through the requirements, a capacity estimate, the two request paths, and the failures that change the design. The numbers below are interview assumptions, not claims about YouTube's actual traffic.
Clarify the scope before drawing boxes
Ask whether the prompt is about video on demand or live streaming. For a 45-minute interview, state that you will design a user-generated video-on-demand service: creators upload a video, the service prepares playable versions, and viewers can start watching from different devices and network speeds.
Keep recommendations, search ranking, comments, advertisements, live broadcasts, and copyright review out of the first design. They are useful follow-ups, but they should not obscure the core data paths.
- Creators can upload large files over an unreliable connection and see processing status.
- Viewers can find a video by its identifier and begin playback quickly.
- Playback adjusts quality as available bandwidth changes.
- A failed upload, worker, storage request, or CDN cache miss has a defined recovery path.
Make one estimate that shapes the architecture
Pick round assumptions and label them. Suppose 10 million viewers each watch 20 minutes per day. That is 200 million viewing-minutes per day, or about 139,000 average concurrent viewers. At an average delivered bitrate of 1.5 Mbps, that is roughly 208 Gbps of average egress. A 3× peak factor gives about 625 Gbps.
The exact answer is less important than the consequence: a single application server or storage origin cannot serve every segment. Playback needs distributed delivery and caching close to viewers. Use the capacity estimator to change the assumptions and see how the estimate moves.
A defensible high-level design
- 1
Upload API
Authenticates the creator, creates an upload session, reserves metadata, and returns short-lived upload URLs. It carries control messages, not gigabytes of video.
- 2
Object storage
Accepts resumable or multipart uploads directly from the client and stores the original file durably.
- 3
Metadata database and job queue
Records the video state and schedules processing. Commit the state change and job publication atomically, using an outbox or equivalent mechanism.
- 4
Transcoding workers
Claim jobs with expiring leases, produce several codec and resolution renditions, validate the outputs, and retry safely after crashes.
- 5
Manifest and segment storage
Stores immutable media segments and a manifest that lists playable renditions. Publish the manifest only after every required object is present.
- 6
Playback API and CDN
The API authorizes playback and returns metadata or a manifest URL. The player fetches segments from a CDN; the origin serves cache misses.
Upload path: move bytes once, recover from interruption
The browser asks the API to start an upload. The API creates a video record in uploading state and returns a scoped session. The browser then transfers chunks directly to object storage. If the connection drops, it asks which byte range arrived and resumes from there. Google's public YouTube Data API documents this resumable upload pattern; it is useful evidence for the interaction, not a description of YouTube's internal storage stack.
When storage confirms the upload, the API records processing and enqueues a job. A worker writes each attempt under a unique output prefix, so a retry cannot overwrite a successful attempt. It validates the renditions and manifest before a conditional database update changes the video toready. That update is the publish point.
Make processing at-least-once and the effects idempotent. A lease lets another worker retry after a crash; a fencing token prevents the old worker from publishing after its lease expired. A reconciler finds abandoned uploads and output objects that no published video references.
Playback path: keep the origin away from the hot path
The player requests a video manifest, chooses a rendition, and fetches short segments over HTTP. It can switch to a lower or higher bitrate as measured throughput and buffer health change. A CDN caches popular immutable segments near viewers; the metadata API remains small and separate from the media bytes.
This cache is a capacity and correctness boundary. A popular video can create a hot object, so cache placement, eviction, and request coalescing affect origin load. Google's published HALP paper describes an eviction algorithm deployed in YouTube's CDN. The authors report an average 9.1% reduction in byte misses during peak, with 1.8% CPU overhead. That is a concrete example of cache policy being a production design problem, not just a box on a diagram.
The exact streaming protocol is a design choice. Apple's HLS documentation describes delivery through ordinary web servers and CDNs, with multiple bitrate streams and switching as network conditions change. In an interview, explain the manifest and segment flow, then name the trade-off between startup delay, segment size, bitrate efficiency, and cacheability.
Failures that expose whether the design is complete
- The client loses its connection mid-upload
- Resume from the acknowledged byte range. Expire abandoned sessions and clean up incomplete multipart data.
- A transcoder crashes after writing some renditions
- Let the lease expire, retry under a new attempt prefix, and publish only after validation. Never infer readiness from a worker heartbeat.
- The queue is healthy but processing falls behind
- Measure oldest-job age and estimated work, not only queue depth. Scale workers against CPU capacity and apply admission limits if backlog threatens the ready-time objective.
- A viral video causes an origin spike
- Cache immutable segments at the edge, collapse concurrent misses where possible, and pre-warm only when popularity is predictable enough to justify the cost.
- A creator deletes a video while it is processing
- Use a terminal deleted state and conditional transitions. Stop or ignore later worker completion, then remove raw and derived objects asynchronously.
- A manifest is visible before one segment is available
- Treat manifest publication as the final atomic step. Keep incomplete attempts private and serve only a manifest whose referenced objects passed validation.
A concise interview answer
“I'll design on-demand uploads and playback first. Clients upload resumably to object storage, while a durable job queue feeds retryable transcoders. A video becomes visible only after all required renditions and its manifest are verified. Playback uses a small metadata API and adaptive segments served through a CDN. I'll size the egress path first, then check worker backlog and storage growth. The main correctness boundary is publication: retries may duplicate compute, but only one validated attempt can become ready.”
Then invite the interviewer to choose a follow-up: live streaming, recommendation feeds, search, copyright review, or the economics of storing every rendition. Do not add those subsystems before the core flows work.
Published sources and scope
- YouTube Data API: Resumable Uploads documents how a client resumes after interrupted video uploads.
- Google Research: HALP for YouTube's CDN describes one published cache-eviction system and its production evaluation.
- Apple Developer: HTTP Live Streaming explains adaptive delivery of multiple bitrate streams over ordinary web infrastructure.
This is an original interview exercise, not an account of private YouTube internals. The public sources above support specific mechanisms; the workload numbers and end-to-end design are stated assumptions for practice.
YouTube system design interview questions
- How do you design YouTube in a system design interview?
- Separate the upload and playback paths. Uploads go directly to durable object storage, then a queue feeds retryable transcoding workers. Publish a video only after its renditions and manifest are verified. Playback reads metadata from a service and streams small adaptive-bitrate segments through a CDN.
- What is the hardest part of a YouTube system design?
- The hardest part is balancing two very different workloads: slow, large uploads and asynchronous processing on one side; latency-sensitive playback at very high read volume on the other. Treating both as ordinary API requests creates timeouts, worker backlogs, and expensive origin traffic.
- Should recommendations be part of the YouTube system design answer?
- Only if the interviewer asks for them. First establish a reliable video upload and playback service. Search, recommendations, comments, ads, live streaming, and creator analytics are separate systems that can each become a follow-up design problem.
- Does this describe YouTube's internal architecture?
- No. It is an interview design exercise with explicit assumptions. The cited Google research describes specific published YouTube systems, but the proposed end-to-end design is not a claim about Google's private production architecture.
The most useful next step is to defend the upload path when a worker dies halfway through a job.