Stage 1 of 14 · Model
Read the requirements like an engineer
Before choosing anything, turn the brief into numbers. Two of them decide most of this design: how long a large upload takes over an ordinary connection, and how many transcodes are in flight at once.
Useful arithmetic: 1 GB is 8 gigabits. Little's law says the average number of jobs in a system equals the arrival rate times the time each job spends there (L = λW). It holds for any stable system regardless of how arrivals are distributed.
Decide whether each statement holds, fails, or depends on something the brief has not settled.
What you need to know first
Network speeds are quoted in bits per second; file sizes in bytes. One byte is 8 bits, so 1 GB is 8 gigabits. To get a transfer time, convert the file to bits and divide by the link speed.
A 4 GB file uploads over a 20 Mbps home connection. About how many minutes does it take?
About 27 minutes.
4 GB × 8 = 32 gigabits = 32,000 megabits. 32,000 ÷ 20 = 1,600 seconds ≈ 27 minutes.
Any server on the path of those bytes has to keep a connection alive for half an hour, through its own timeouts and deploys.
Little's law: the average number of jobs in a system equals the arrival rate times the time each job spends there.
L = λ × WIt holds for any stable system, whatever the pattern of arrivals. It turns "how many per second" and "how long each" into "how many at once", which is the number that sizes a worker fleet.
Sunday peak: about 0.12 uploads arrive per second, and each transcode takes 750 seconds. How many transcodes are in flight at once if nothing waits?
About 90 transcodes.
L = 0.12 × 750 = 90 transcodes in flight.
With 20 workers, the other 70 wait in a queue, and the queue grows for as long as the peak lasts.
A worker dies 14 minutes into a 20-minute transcode. The requirement says accepted uploads are never lost. What does that force?
The job must be able to run again, so running twice must be harmless.
The only alternative to re-running is never finishing, which breaks the requirement. And since a slow worker looks like a dead one, sometimes both will run.
What the stage asks
Which of these follow from the requirements and constraints?
- Holds
A 4 GB upload over a 20 Mbps home connection takes roughly half an hour.
4 GB is 32 gigabits; at 20 megabits per second that is 1,600 seconds, about 27 minutes. A 10 GB file takes over an hour. Any component on the path of those bytes has to tolerate a connection that lives this long.
- Fails
Uploads can pass through the API servers as long as they have enough memory and disk.
Capacity is not the issue. The load balancer cuts requests at 60 seconds, and the containers are redeployed several times a day, so a 27-minute upload through them would be cut off by the timeout and again by every deploy. Raising the timeout fixes the first problem but not the second: a redeploy still kills every in-flight upload.
- Holds
Transcoding has to happen outside the request that delivers the upload.
A 5-20 minute transcode cannot finish inside a 60-second request, and even with a longer timeout, tying the response to minutes of CPU means a deploy or crash loses the work and the client has no way to find out. The work has to outlive the request: see Asynchronous processing.
- Depends
Twenty always-on workers are enough for launch traffic.
Over a whole day: 2,000 × 12.5 minutes ≈ 17 worker-days of CPU, so twenty workers keep up on average. During the Sunday peak the arrival rate is 5x, about 0.12 uploads per second, and Little's law gives 0.12 × 750 s ≈ 87 transcodes in flight if nothing waits. With twenty workers a backlog builds all evening and takes hours to drain.
Whether that is acceptable is a product decision: how long may an instructor wait on a Sunday? The queue turns a capacity problem into a latency problem; it does not make it disappear. See Backpressure and capacity.
- Holds
Because workers can crash mid-job, the design must expect some jobs to run more than once.
If a worker dies at minute 14, either the job runs again or it never finishes. "Never finishes" violates a requirement, so re-execution is the only option, and from the outside a slow worker is indistinguishable from a dead one. The goal is not exactly-once execution, which nobody can promise here, but making a second execution harmless.
The reasoning
- Convert sizes to bits and divide by link speed: a 4 GB upload over home broadband takes about half an hour.
- Little's law (L = λW) turns arrival rate and job time into concurrent jobs, which sizes the fleet.
- When workers can crash, jobs must be safe to run twice.
The two numbers to keep are ~30 minutes for an upload and ~90 concurrent transcodes at peak. Together they rule out most of the prototype.
- The upload outlives every timeout and deploy cycle on the API tier, so the bytes cannot flow through it.
- The transcode outlives any reasonable request, so it must become a durable job that some other process picks up.
- Once work is a job that survives crashes, it can run twice, and the rest of the design has to make that safe.
Notice that the capacity question has no single correct answer. It turns into a choice between paying for peak capacity and making people wait. A design review that skips that conversation has decided it by accident.
Tradeoffs
| Choice | Gains | Costs |
|---|---|---|
| Provision for the Sunday peak (~90 workers) | Videos start processing immediately at all times. | Most of that capacity sits idle six and a half days a week. |
| Provision near the daily average and let a backlog form | Roughly a quarter of the compute cost. | Sunday-evening uploads can wait hours, so the status UI has to say so honestly. |