The three binaries that matter
A build produces several executables. Three do almost all the work.
llama-cli runs inference from the command line — one-shot completion or an interactive session. It is where you test whether a model and a set of parameters do what you want.
llama-server exposes an HTTP API and a minimal web interface. The API implements the OpenAI chat-completions shape, so applications written against that specification can be pointed at it by changing a base URL. This is how llama.cpp ends up inside other software.
llama-quantize converts a model from one numeric format to a smaller one. It is the tool that makes everything else possible on consumer hardware.
Names and flags have changed across releases — older tutorials refer to main and server rather than the current llama- prefixed binaries. When a command from a guide does not exist, this is usually why.
Quantisation, and what each level actually costs
Model weights are normally 16-bit floating point. Quantisation stores them in fewer bits, trading accuracy for memory. This is the single most important concept for anyone running models locally, and the naming is opaque until explained.
A file marked Q4_K_M decodes as: 4 bits per weight, K-quant method, Medium size variant. The K-quant methods allocate bits unevenly — more precision to the layers that matter most — which is why a modern 4-bit model is far better than the number suggests.
Roughly what the levels mean in practice:
| Q8_0 | Near-indistinguishable from full precision. Rarely worth the memory. |
| Q6_K | Very close to the original. A safe choice when memory allows. |
| Q5_K_M | Small, consistent degradation. A common sweet spot. |
| Q4_K_M | The usual default. Noticeable on hard reasoning, fine for most work. |
| Q3_K | Degradation becomes visible — repetition, weaker instruction following. |
| Q2_K | Usually a false economy. Run a smaller model at Q4 instead. |
There is no warning when a model is degraded. Output simply gets worse, and the failure is easy to blame on the model rather than the quantisation. A useful rule: prefer a smaller model at higher precision over a larger model crushed to 2 or 3 bits.
Working out whether a model will fit
Two things consume memory, and people usually account for only the first.
The weights. Approximately: parameters × bits-per-weight ÷ 8. A 7-billion-parameter model at 4 bits is roughly 3.5 GB, plus a little overhead.
The KV cache. Every token in the context window is held in memory during generation, and this scales with context length. A long context can consume as much memory as the model itself. Setting a 32k context on a machine that barely fits the weights is the most common cause of an out-of-memory failure part-way through a session rather than at load.
If the total exceeds available memory, --n-gpu-layers splits the model: the specified number of layers move to the GPU and the remainder stay on the CPU. A card with insufficient VRAM still contributes rather than being unusable, and finding the highest number that loads is normal practice.
Building it, and where builds usually fail
- A C/C++ toolchain and CMake. Prebuilt binaries are published per release, but building from source is how you enable a specific accelerator.
- Accelerator support is a build-time decision, not a runtime flag. CUDA for NVIDIA, Metal on Apple silicon, Vulkan, SYCL for Intel, ROCm for AMD — each is a separate build configuration.
- Models in GGUF format. Most open-weights models are published or converted to it; a model available only in another format needs converting first.
- Enough RAM for weights plus KV cache, as above.
The recurring build failure is a CUDA toolkit that does not match the installed driver, which produces a compile error rather than a helpful message. On Apple silicon, Metal support is enabled by default and rarely causes trouble. No GPU is required at all — CPU inference is a supported first-class path, not a fallback.
What it deliberately will not do for you
llama.cpp is an inference engine. It does not train or fine-tune models, does not manage a model library, does not download anything, and provides no meaningful interface beyond a minimal server page.
It also makes no decisions. Model, quantisation, context size, thread count, batch size and layer offload are all explicit arguments with no sensible defaults chosen for your hardware. Every wrapper in this category exists to make those choices on your behalf.
Nothing is sent anywhere. There is no account, no telemetry and no network dependency once the model is on disk — which is the entire reason it appears in air-gapped and regulated deployments.
Where it earns its place over a wrapper
Developers embedding local inference into an application, where a wrapper’s assumptions get in the way and a single compiled binary is easier to ship than a runtime.
Anyone tuning for specific hardware — fitting a model into a particular VRAM budget, or extracting acceptable speed from an older CPU. The controls that matter are exposed here and hidden elsewhere.
Packagers and researchers who need reproducible, dependency-light builds, and deployments to constrained targets such as ARM boards or offline machines where a Python stack is impractical.
And anyone who wants the newest model architectures. Support generally lands here before it reaches the tools built on top.
It is the wrong starting point for someone who simply wants to chat with a local model. That is what the wrappers are for, and using them is not a compromise.
Licensing, and why it matters commercially
llama.cpp is MIT licensed. You can embed it in a commercial product, modify it, and ship it without publishing your changes or entering a licensing conversation.
That is unusually permissive for infrastructure of this importance, and it is why it appears inside so much commercial software.
The engine’s licence is not the model’s licence. Whether you may use a model’s output commercially depends entirely on the weights you load — see Stable Diffusion and FLUX for how sharply those terms differ within a single family. Check the model, every time.
What you gain by using it directly
- Control of the accuracy/memory trade-off — you choose the quantisation rather than accepting a default.
- Almost no dependencies, which makes it practical to embed and to audit.
- Broad hardware support: Apple silicon, NVIDIA, AMD, Intel and plain CPU are all supported paths.
- Layer-level GPU offload, so partial acceleration is possible rather than all-or-nothing.
- An OpenAI-compatible server without installing a separate component.
- MIT licensed, so commercial embedding needs no permission.
- First to support new architectures, usually by a wide margin.
The costs, stated honestly
- You have to build it. Getting GPU acceleration working is where most people give up, and the errors are compiler errors rather than guidance.
- Quantisation degrades output silently. No warning, no indicator — just worse answers that look like a bad model.
- Command line only. The bundled web UI is deliberately minimal.
- The interface moves. Binary names and flags have changed across releases, so older documentation frequently does not apply.
- Single-machine inference. For concurrent users, vLLM is the correct tool and this is not close.
- No model management. Finding, downloading and organising GGUF files is your problem.
If you leave, what do you take with you
Almost everything. Models are GGUF files on your disk and are usable by every other tool in this family — Ollama, LM Studio and Jan all read the same format. Prompts and scripts are yours.
There is no account to close, no data held elsewhere, and nothing that stops working if the project does. That is a materially different position from every hosted tool in this directory, and it is worth weighing against the setup cost.
Wrappers worth using instead
- Ollama — llama.cpp with model management and a one-line install. The sensible default unless you need the control, and the one we have hands-on tested.
- LM Studio — a desktop application over the same engine, with a model browser that tells you what will fit.
- vLLM — GPU-oriented serving for concurrent requests. A different problem entirely.
- LocalAI — wraps llama.cpp among other backends behind one OpenAI-compatible API.
Compiled from the llama.cpp project documentation and public sources. We have not hands-on tested this tool. Last reviewed 16 August 2026.