LLM inference capacity calculator
Estimate model-weight memory, KV-cache memory for active sequences, and the input and output token rates your serving design must handle. Use it to make a system design interview estimate, then benchmark the model and serving runtime on the target hardware before making a production capacity plan.
Free · In-browser · No account
Estimate LLM GPU memory and token workload
How the LLM inference estimate works
Weight memory starts with parameter count multiplied by bytes per parameter. The KV cache stores keys and values for each active sequence. For a uniform full-attention model, the estimate per token is 2 × layers × KV heads × head dimension × bytes per cache value. The factor of two accounts for keys and values. Multiply that by the prompt plus reserved output length and then by active sequences. The calculator uses the model's key/value-head count, which is important for grouped-query attention.
The estimate adds model weights, a runtime/workspace reserve you choose, and KV cache for your target concurrency. It compares that sum with the selected fraction of aggregate GPU memory. Weight and KV memory can be partitioned differently by a serving runtime, so an aggregate fit is only a first check. NVIDIA's TensorRT-LLM memory guide describes weights, activation tensors, and KV cache as distinct inference memory contributors.
Prompt tokens per second and generated tokens per second are separate workload estimates: request rate multiplied by average prompt length and output length, respectively. They are inputs for a benchmark plan, not a predicted GPU throughput. The token-rate estimates use averages; the KV-cache estimate uses the prompt and output budgets for active sequences. Measure time to first token, inter-token latency, throughput, and memory use with the target model, runtime, and length distribution. The Hugging Face KV-cache guide explains how cache size grows with sequence length and model architecture.
Use the estimate in an interview answer
- State the model, precision, prompt/output length distribution, concurrency, and latency target.
- Estimate input tokens per second, output tokens per second, and KV cache for the active sequences.
- Compare the memory estimate with one replica's GPU pool; include a reserve for runtime memory and fragmentation.
- Explain what you would benchmark and what your scheduler does when the queue or KV cache is full.
LLM inference memory and capacity questions
- How do you estimate KV-cache memory for an LLM?
- For a full-attention decoder with uniform layers, estimate bytes per token as 2 × layers × key/value heads × head dimension × bytes per cache value. Multiply by the prompt plus reserved output tokens per active sequence, then by active sequences. Grouped-query attention uses the number of key/value heads, not the query-head count.
- Can this calculator tell me how many GPUs an LLM needs?
- No. It compares a rough model-weight and KV-cache estimate with the aggregate memory in one model replica. It does not predict accelerator throughput, tensor-parallel efficiency, runtime workspaces, or whether the model is evenly shardable. Benchmark the chosen model and serving stack on the target GPUs before planning production capacity.
- Why estimate input and output tokens per second separately?
- Prompt processing and autoregressive generation have different execution behavior and bottlenecks. Multiply requests per second by average prompt and output lengths for token-demand estimates. For KV-cache capacity, use the prompt and output budgets you intend each active sequence to hold.
- Is this SysGeeks LLM inference calculator free?
- Yes. It runs in your browser, requires no account, and does not send the values you enter to the site.