PagedAttention, and the memory problem it solves
Understanding this one idea explains why vLLM exists and when it is worth the operational burden.
During generation, a model holds a KV cache — the attention state for every token so far — in GPU memory. Conventional serving reserves a contiguous block per request, sized for the longest output that request might produce. A request that could generate 2,000 tokens but stops at 100 has reserved memory for 2,000 and wasted most of it.
At scale this is the binding constraint. Reported waste in naive implementations runs to the majority of allocated cache memory, which caps concurrency far below what the hardware could support.
PagedAttention manages the cache in fixed-size blocks, the way an operating system manages virtual memory with pages. Memory is allocated as tokens are actually produced, and blocks need not be contiguous. Waste falls to a few percent, and the same card serves far more simultaneous requests.
It also enables prefix sharing: several requests beginning with the same system prompt can share those blocks rather than each holding a copy — a substantial saving in any application with a long fixed preamble.
Continuous batching, and why latency stays reasonable
The second idea is about scheduling. Static batching groups requests, runs them together, and returns when the slowest finishes — so a short request waits on a long one and the GPU idles at the end of every batch.
vLLM schedules at the token level. Completed sequences leave the batch immediately and queued requests join mid-flight, so the GPU stays saturated and short requests are not held hostage by long ones.
Together these two mechanisms are the whole argument for vLLM. Neither matters at all if you are the only user.
Hardware, and the one number that decides everything
- An NVIDIA GPU with CUDA is the primary supported path. AMD ROCm, Intel, AWS Neuron, TPU and CPU backends exist; CUDA is where the project’s attention goes and where problems get fixed first.
- VRAM is the binding constraint. The model weights must fit, plus KV cache for the concurrency you intend to serve. A model that fits with no cache headroom will serve one user badly.
- Linux and Python. Installation is pip or a container image.
- Models in Hugging Face format — read directly, with no conversion step, unlike the GGUF ecosystem.
For models too large for one card, tensor parallelism splits each layer across GPUs and pipeline parallelism splits the layers between them. Tensor parallelism needs fast interconnect between cards to be worth it; over a slow link the communication cost eats the benefit.
Quantised weights (AWQ, GPTQ and similar) reduce the memory requirement at some cost in quality, which is often what makes a given model fit a given card at all.
The API surface, and what migrating actually involves
vLLM serves an HTTP endpoint implementing the OpenAI API — /v1/chat/completions, /v1/completions and /v1/embeddings. Existing client code is repointed by changing a base URL and supplying any non-empty key.
That makes migration mechanically trivial and semantically not. The API is compatible; the model is not the same model. Prompts tuned against a frontier hosted model frequently need revisiting, output formats drift, and behaviour under edge cases differs. Budget for prompt work, not for integration work.
Less common parameters may be unsupported or silently ignored, so an application leaning on the edges of the specification needs testing rather than assuming.
What it is not, and should not be asked to be
vLLM serves models. It does not train or fine-tune them, provides no user interface, and manages no model catalogue. It is infrastructure that sits behind something else — an application, an agent framework, or a front-end such as Open WebUI.
It is also not a desktop tool. Running it on CPU is possible and useful for development, and gives up the entire reason to use it.
Who this is genuinely for
Teams putting an open-weights model into production: a product feature, an internal assistant serving a whole company, or a batch pipeline processing large volumes where per-token pricing would be prohibitive.
Organisations with data residency or privacy requirements that rule out a hosted API, but who still need the concurrency a hosted API provides.
And anyone who has done the arithmetic. Roughly: below a few million tokens a month, an API is cheaper once staff time is counted. Above that, and especially with predictable sustained load, owned hardware wins. The crossover depends on your GPU cost and your engineers’ time, and it is worth calculating rather than assuming.
It is the wrong tool for a single developer wanting a local coding assistant — Ollama does that with a fraction of the effort.
Running it in production, which is the real cost
The install is easy. Operating it is not, and this is the part most evaluations underestimate.
Someone has to monitor GPU memory and request latency, size the cache against expected concurrency, handle restarts when a bad request wedges the server, manage model updates, and be available when it fails at an inconvenient hour. A hosted API includes all of that in the price.
Realistically this needs a person who is comfortable with CUDA drivers, container orchestration and GPU monitoring. If nobody on the team owns that, the project stalls after the proof of concept.
Its genuine advantages
- Substantially more concurrent requests per GPU than naive serving — the reason it exists.
- Prefix sharing, which is a large saving for applications with long system prompts.
- Drop-in OpenAI-compatible endpoint, so migration is a base-URL change.
- Multi-GPU support for models too large for one card.
- Predictable cost — hardware you already run, rather than billing that scales with success.
- Apache 2.0 licensed, with no commercial restrictions.
- Fast support for new architectures, and a large active contributor base.
Where it will frustrate you
- GPU-bound. Without a suitable NVIDIA card, most of the value is unavailable.
- Real operational burden, which is the honest cost and is rarely in the evaluation.
- Not worth it below a certain volume — at low request rates hardware plus operations exceeds what an API would have charged.
- Linux-first. Windows via WSL or containers rather than native.
- Model support varies. Broad coverage of common architectures; a brand-new or unusual model may need work.
- Cold starts are slow. Loading a large model into GPU memory takes time, so scaling to zero is impractical.
Leaving, and what it costs you
Little. Models are standard Hugging Face repositories, and your application talks the OpenAI API — so moving to a hosted provider, or to another self-hosted server, is a configuration change rather than a rewrite.
That is a deliberate property of choosing an API-compatible server, and it is worth preserving: keep provider-specific behaviour out of your application code and the exit stays cheap.
Adjacent tools
- Ollama — single-machine local inference, far simpler, no concurrency story.
- llama.cpp — the CPU-capable engine under most desktop tools. A different problem.
- LocalAI — OpenAI-compatible API over lighter backends, for smaller deployments.
- Open WebUI — pairs with vLLM to give a served model a usable interface.
Compiled from the vLLM project documentation and public sources. We have not hands-on tested this tool. Last reviewed 16 August 2026.