Skip to content

A reliable video processing pipeline

Design a Video Processing Pipeline

Instructors upload multi-gigabyte lectures that take minutes to transcode. Workers crash, deploys interrupt jobs, and the same job can run twice. Every accepted upload must end in exactly one correct, visible outcome.

Intermediate, about 55 minutes, 14 stages

The situation

A course platform lets instructors upload lecture recordings. Files are usually 200 MB to 4 GB, recorded on laptops and uploaded over home or café connections. Each upload has to become an adaptive-bitrate stream (1080p, 720p and 480p renditions) plus a thumbnail before students can watch it.

The prototype is an Express server that accepts a multipart form upload to local disk and runs ffmpeg inside the request handler. It worked for the ten test videos. The platform launches next month at around 2,000 uploads a day, with Sunday evenings running at five times the average rate.

You own the pipeline from "instructor selects a file" to "student presses play". Nothing about the prototype is sacred.

What it has to do

Functional

  • Instructors upload video files of up to 10 GB.
  • Each upload becomes a 1080p/720p/480p HLS ladder and a thumbnail.
  • Instructors see each video's status: uploading, processing, ready, or failed with a reason.
  • Students can stream a video only once every rendition is ready.
  • Instructors can delete a video at any time, including while it is processing.

Non-functional

  • An upload the platform has accepted is never silently lost.
  • A video is never shown as ready unless every rendition exists and plays.
  • A crashed or redeployed worker must not strand a job forever.
  • Duplicate processing may waste compute but must never produce duplicate or corrupt output.
  • API servers stay responsive during processing peaks.

Constraints and assumptions

  • About 2,000 uploads a day at launch, peaking at 5x the average on Sunday evenings; 10x growth expected within a year.
  • A transcode takes 5-20 minutes of CPU (about 12.5 on average); one worker runs one transcode at a time.
  • API servers are stateless containers behind a load balancer with a 60-second request timeout, redeployed several times a day.
  • Small team: managed infrastructure is preferred, and one Postgres database already exists.
  • An object store with S3-like semantics is available: durable writes, presigned URLs, multipart uploads, and read-after-write consistency for new objects.
  • Instructors' browsers can upload to the object store directly over HTTPS.
  • Transcoding the same input twice produces equivalent output.
  • Postgres is the system of record; losing it is a disaster-recovery event, not a design input.

Interview questions it prepares you for

  • “Design YouTube's upload and processing pipeline.”
  • “Design a service that generates large PDF reports that take minutes to produce.”
  • “Design a background job system that survives worker crashes and deploys.”
  • “How would you generate thumbnails for millions of image uploads a day?”
  • “Your job queue sometimes runs the same job twice. What do you do?”

Read and practise next

How Mux built it · Video that plays before it is processed, in their engineers' own words

Concepts to know first: Asynchronous processing, Object storage.

Similar systems: Design a Payment System, Design a Collaborative Editor (Google Docs).