Skip to content

Design a Video Processing Pipeline, stage 12 of 14: change it

100x traffic: find the real bottleneck

Before you change anything, work out what actually breaks. "It needs to scale" is not a diagnosis. Find the resource that runs out first.

System so far· 8 parts
123456789CLIENTInstructorbrowserSERVICEVideo APIDATABASEPostgresOBJECT STOREObject storageWORKERTranscodeworkersWORKERReconcilerEDGECDNCLIENTStudent player

Select a component to see what it is responsible for and which state it owns.

  1. 1Instructor browser → Video API: Create upload, report parts, poll status
  2. 2Instructor browser → Object storage: Upload parts via presigned URLs
  3. 3Video API → Object storage: Complete multipart upload, verify object
  4. 4Video API → Postgres: Video row and job row in one transaction
  5. 5Transcode workers → Postgres: Claim lease, heartbeat, fenced completion
  6. 6Transcode workers → Object storage: Read raw upload, write attempt output
  7. 7Reconciler → Postgres: Find abandoned uploads and orphaned output
  8. 8CDN → Object storage: Origin fetch on cache miss
  9. 9Student player → CDN: Manifest and segments
  • Request / response
  • Bulk data

What you need to know

0 of 3 checks done
  1. "It needs to scale" isn't a diagnosis. For each resource the system uses (coordination writes, compute, storage, network out), compute the demand at the new volume and compare it with what that resource can supply. The first one to run out is the bottleneck. Fixing anything else changes nothing.

  2. Work it out

    200,000 uploads a day. About how many jobs a second is that at the 5× Sunday peak?
    per second