Design a Video Processing Pipeline, stage 9 of 14: break it
Write the claim and the completion
Turn the last two stages into code. You need two operations: one that atomically claims the next available job, and one that completes a job only if the caller still owns it. SQL, an ORM, or pseudo-code are all fine. What matters is which conditions are checked, and where.
System so far· 8 parts
Select a component to see what it is responsible for and which state it owns.
- 1Instructor browser → Video API: Create upload, report parts, poll status
- 2Instructor browser → Object storage: Upload parts via presigned URLs
- 3Video API → Object storage: Complete multipart upload, verify object
- 4Video API → Postgres: Video row and job row in one transaction
- 5Transcode workers → Postgres: Claim lease, heartbeat, fenced completion
- 6Transcode workers → Object storage: Read raw upload, write attempt output
- 7Reconciler → Postgres: Find abandoned uploads and orphaned output
- 8CDN → Object storage: Origin fetch on cache miss
- 9Student player → CDN: Manifest and segments
- Request / response
- Bulk data
What you need to know
A claim has to be atomic: choosing a job and marking it taken must be one step, or two workers can choose the same job. In Postgres, one
UPDATE … WHERE id = (SELECT … FOR UPDATE SKIP LOCKED LIMIT 1)statement does both.The row lock lasts only for that statement. The lease is what holds ownership for the 20 minutes the transcode takes.
Check
Where should the attempt counter be incremented?Completion makes two changes that must happen together: the job becomes
succeeded, and the video becomesreadywith its manifest pointer. Both go in one transaction, and the job update carries the lease token. If the token no longer matches, the whole completion is abandoned and this attempt's output is left unpublished.Think first
A worker's heartbeat update changes zero rows. What should the worker do?