Local Inference Lab

How much LLM can one 16 GB GPU serve?

Real benchmarks of vLLM, llama.cpp and Ollama serving the same 7B model at 4-bit, and a KV cache planner you can run below. Loading results…

Which server, and for how many users

Each config served the same requests at rising concurrency. Throughput is output tokens per second across all users; time to first token is how long a user waits before the answer starts.

Throughput (tokens/s, higher is better)
Median time to first token (seconds, log scale, lower is better)

KV cache: FP16 vs FP8

vLLM turns the memory left after the weights into KV cache, which holds every active user's context. Storing it in FP8 halves the bytes per token, so more users fit before vLLM has to pause (preempt) requests.

Plan a vLLM deployment

Runs in your browser with the same formula as inference-lab plan vllm. Weights size is what the checkpoint's files add up to; overhead covers activations and CUDA graphs.