Skip to content

Design a Video Processing Pipeline, stage 10 of 14: break it

The video that kills every worker

The lease mechanism faithfully recovers the crashed job, again and again. Meanwhile other instructors' videos wait behind a job that will never succeed, and this instructor sees "processing" indefinitely.

System so far· 8 parts
123456789CLIENTInstructorbrowserSERVICEVideo APIDATABASEPostgresOBJECT STOREObject storageWORKERTranscodeworkersWORKERReconcilerEDGECDNCLIENTStudent player

Select a component to see what it is responsible for and which state it owns.

  1. 1Instructor browser → Video API: Create upload, report parts, poll status
  2. 2Instructor browser → Object storage: Upload parts via presigned URLs
  3. 3Video API → Object storage: Complete multipart upload, verify object
  4. 4Video API → Postgres: Video row and job row in one transaction
  5. 5Transcode workers → Postgres: Claim lease, heartbeat, fenced completion
  6. 6Transcode workers → Object storage: Read raw upload, write attempt output
  7. 7Reconciler → Postgres: Find abandoned uploads and orphaned output
  8. 8CDN → Object storage: Origin fetch on cache miss
  9. 9Student player → CDN: Manifest and segments
  • Request / response
  • Bulk data

What you need to know

0 of 2 checks done
  1. Not every failure is worth retrying. Retries help only when the next attempt might behave differently:

    KindExampleRetry?
    Transientstorage returned 503, instance reclaimedyes, with backoff
    Permanentnot a valid video, unsupported codecno: fail now, with a reason
    Unknownthe worker diedyes, but within a budget
  2. Check

    A corrupt file crashes ffmpeg 40 seconds into every attempt. Leases recover the job each time. With no attempt limit, what happens?