How much LLM can one 16 GB GPU serve?
Real benchmarks of vLLM, llama.cpp and Ollama serving the same 7B model at 4-bit, and a KV cache planner you can run below. Loading results…
Which server, and for how many users
Each config served the same requests at rising concurrency. Throughput is output tokens per second across all users; time to first token is how long a user waits before the answer starts.
KV cache: FP16 vs FP8
vLLM turns the memory left after the weights into KV cache, which holds every active user's context. Storing it in FP8 halves the bytes per token, so more users fit before vLLM has to pause (preempt) requests.
Plan a vLLM deployment
Runs in your browser with the same formula as inference-lab plan vllm. Weights size is
what the checkpoint's files add up to; overhead covers activations and CUDA graphs.