vLLM is an inference and serving engine built for throughput rather than single-user convenience. It uses paged attention and continuous batching to serve many concurrent requests efficiently on GPU hardware, and exposes an OpenAI-compatible API so existing clients work unchanged. It is the layer you reach for when a local model needs to serve an application or a team rather than one person at a terminal.
Pricing: Free