Skip to content

Design an LLM inference service: a system design interview walkthrough

Design an API that serves text from self-hosted language models to multiple customers. The key decisions are how to admit work before GPUs are saturated, how to trade batching against response latency, and what to do when a worker fails mid-stream. This is an interview design with explicit boundaries, not a claim about any provider's private architecture.

Clarify the system before choosing GPUs

Start by asking whether the service runs its own model weights or routes to an external model provider. This walkthrough assumes self-hosted, open-weight models, a text completion API, multiple customer accounts, and streamed output. Model training, retrieval, tools, and multimodal inputs are separate extensions; add them only if the interviewer includes them.

  • Customers authenticate, choose an allowed model, set generation limits, and receive a token stream.
  • Each request uses one model version for its full lifetime, including after a rollout begins.
  • Usage is attributed to a customer so quotas and invoices can be reconciled.
  • Overload, cancellation, timeouts, and worker failures have clear behavior for the caller.

Ask for the target time to first token, acceptable delay between streamed chunks, request volume, prompt and output lengths, availability, and cost constraints. If those numbers are not given, state assumptions and label them before estimating.

Estimate token throughput alongside request rate

For a request rate R, average prompt length P, and average generated length G, estimate input work as R × P prompt tokens per second and output work as R × G generated tokens per second. Keep the two figures separate: prompt processing and autoregressive generation stress the serving path differently.

Request QPS hides the difference between short and long prompts, answers, and concurrent generations. Ask for their distributions as well as averages; a small tail of very long contexts can occupy memory and delay other requests. The capacity estimator helps with rough request rate and network volume, but it does not predict GPU count. Benchmark the actual model, accelerator, context mix, and latency target before making a hardware estimate. The official vLLM benchmark guide describes serving measurements such as time to first token and inter-token latency.

A defensible high-level architecture

  1. 1

    API gateway

    Terminates TLS, authenticates the API key, applies request-size limits, and attaches the customer identity. It rejects malformed or unauthorized calls before they can consume accelerator time.

  2. 2

    Admission controller

    Checks model access, context length, max output tokens, tenant quota, and current capacity. It reserves a bounded amount of work rather than treating every request as one equal-sized unit.

  3. 3

    Model router and scheduler

    Routes each admitted request to a compatible model-version pool. A bounded queue applies deadlines and fair scheduling; when the queue is full, return an overload response instead of allowing wait time to grow without limit.

  4. 4

    GPU inference workers

    Keep model weights loaded and schedule compatible sequences for prefill and token generation. Use the runtime's batching and KV-cache management, and expose queue, memory, and token-throughput signals for capacity control.

  5. 5

    Streaming response

    Returns generated chunks as they are ready, with a request identifier and a clear terminal status. A disconnected client cancels queued work or asks the worker to stop generation when possible.

  6. 6

    Usage ledger

    Emits an idempotent usage event keyed by request and attempt. A separate consumer aggregates token counts for quotas and billing so a billing-store slowdown does not block token delivery.

Store model manifests and immutable weight artifacts in a registry and object store. Workers load a pinned version before becoming healthy; the router sends new requests to a canary pool first, while in-flight streams stay on the version that accepted them.

Balance batching throughput against queue time

Continuous batching lets the scheduler admit new requests and remove completed ones as generation proceeds, so a batch does not have to wait for its slowest response to finish. A runtime may also group stateless inference calls dynamically. Both approaches can raise accelerator utilization, but a batching delay adds time before work starts. Tune batch size and delay against the latency budget, and measure the result on the chosen workload. See the official vLLM serving documentation and NVIDIA Triton batcher guide.

Bound active sequences and queued tokens per model pool. A request with a long prompt or a large maximum output budget can consume much more memory and time than a short request. Use per-customer concurrency limits and a fair scheduler so one tenant cannot monopolize the GPU queue. If the system cannot meet a deadline, reject early with a retryable overload response or route to a compatible fallback model when the product contract allows it.

Keep prompt caching optional and scoped to data the customer is allowed to reuse. A cache hit can skip repeated prompt work, but it does not remove the need to check tenant boundaries, model version, and freshness. Avoid presenting a cache as a substitute for token and queue limits.

Trace the failures that change the design

The queue grows during a traffic spike
Queue time becomes part of the first-token delay. Bound queued requests and prompt tokens, shed load before the latency target is impossible, and scale on queue time and token throughput rather than request count alone.
A long context fills the KV cache
Reject requests over the model's context limit, bound active token budgets, and watch cache usage and preemptions. A larger queue does not create more GPU memory.
A worker dies before the first token
If the request is known not to have started, another healthy replica can accept the same request id under an explicit retry policy. Record the attempt so usage is not counted twice.
A worker dies after some tokens were sent
The client has already observed a partial answer. End the stream clearly; do not silently restart and concatenate a second generation. Preserve the request and attempt ids so support and billing can explain the outcome.
A new model version behaves badly
Stop routing new requests to the canary and shift traffic back. Do not move an active stream between versions; keep weights immutable and track which version handled each request.
The client disconnects
Cancel work waiting in the queue immediately. Propagate cancellation to generation where supported, then record whether any tokens were produced so usage follows the stated billing rule.

Measure the user-visible bottleneck

Track request queue time, time to first token, inter-token latency, end-to-end latency, prompt and generated tokens, active and waiting requests, KV-cache usage, preemptions, cancellations, and errors. Break the measurements down by model version and workload class; sample customer identifiers carefully so metrics do not expose prompt content.

vLLM documents metrics for queue time, time to first token, inter-token latency, token counts, and KV-cache use in its production metrics guide. Those separate signals help tell whether a slowdown comes from admission, prompt processing, generation, or memory pressure. Watch p95 and p99 alongside the average, since a good average can hide a poor tail.

A concise interview answer

I would first pin down the model, prompt and output length distributions, latency targets, and whether we host the weights. I would estimate prompt and generated tokens per second separately, then admit requests through bounded per-model queues with tenant limits. GPU workers keep a pinned model version loaded and batch compatible sequences; the API streams tokens and propagates cancellation. If the queue or token budget is full, I reject early rather than promise a latency the system cannot meet. I would monitor queue time, first-token time, inter-token latency, token throughput, and KV-cache pressure, then use benchmark results to decide when to add replicas or change the model.

A strong answer connects the request budget, scheduling policy, latency promise, and retry semantics.

Sources and scope

  • vLLM production metrics documents serving metrics, including waiting requests, queue time, token latency, and KV-cache usage.
  • vLLM benchmark CLI describes serving benchmarks for request latency and token throughput.
  • NVIDIA Triton batchers describes dynamic batching, queue limits, timeouts, and the latency-throughput tradeoff.
  • Anthropic Engineering: Managed Agents explains time to first token as a user-visible delay in one production system. That account supports the metric definition; this page's design remains an independent interview exercise.

LLM inference system design questions

How do you design an LLM inference service in a system design interview?
Separate the API and admission path from GPU execution. Authenticate each tenant, validate and bound token usage, route to a pinned model version, admit requests to a bounded scheduler, batch active sequences on GPU workers, and stream results while recording token usage. Protect latency with per-tenant limits, deadlines, and explicit overload behavior.
Why is request QPS not enough to size an LLM service?
Requests have different prompt and output lengths. Estimate input tokens per second as request rate times average prompt tokens, and generated tokens per second as request rate times average output tokens. Then benchmark the chosen model, hardware, context lengths, and concurrency; those measurements determine replica capacity.
What does continuous batching improve?
A continuous batch can admit new requests and remove completed ones as generation proceeds, keeping compatible work on the accelerator. Higher batch occupancy can improve throughput, while queueing and batch delay can worsen time to first token or gaps between streamed tokens. Tune both against the service's latency target.
What happens if a GPU worker fails after streaming has started?
The service cannot safely pretend that a restarted generation is the same response: it may repeat or change tokens. End the stream with an error and request identifier, account for already generated usage according to the product contract, and let the client explicitly retry as a new attempt. A retry before generation begins is a different case.
Can I practise this design on SysGeeks?
Yes. Review the worked answer, then use the model inference API example to inspect a related architecture and the timed mock interview to practise explaining your decisions.

Now defend the queue limit when the interviewer doubles prompt length and keeps request QPS unchanged.