Ollama

Run open-weights AI models locally on your own machine

★★★½☆ 3.7 / 5 How we rate
Rating reviewed 2 Oct 2026 ✓ Hands-on tested
Ollama logo
Pricing Free
Category 💻 AI Coding & Dev
Our Rating 3.7 / 5
Best For Whom Running open-weights models locally without a…

Ollama is the most direct way to run large language models on your own computer. You install it, pull a model with a single command, and talk to it from a terminal or through its OpenAI-compatible API. Nothing is sent to a server and no account is required. It runs on Windows, macOS and Linux, and works on CPU alone when no GPU is available. Its library covers most major open-weights families. A paid cloud tier exists for hosted models, which the launcher surfaces alongside local ones.

Ollama runs large language models on your own machine. You install it, pull a model, and talk to it from a terminal or through its API — no account, no per-token billing, and nothing leaving your computer.

We installed it on a mid-range desktop with no graphics card and measured what actually happens.

✅ Pros

  • Runs on ordinary hardware, no GPU required
  • One-command install and model pull
  • Free and open source, nothing leaves your machine
  • OpenAI-compatible API for local apps

❌ Cons

  • CPU-only generation is slow
  • Model tags are not guessable
  • Launcher promotes paid cloud models
  • Models are large on disk

🎯 Best For What

Running open-weights models locally without a GPU

How we scored Ollama

Ten dimensions, each out of 5. Nine are editorial; the tenth, Demand, is calculated from how often this page is actually read and is refreshed weekly. Full methodology

  • Capability 4/5
  • Ease of use 4/5
  • Value 5/5
  • Reliability 4/5
  • Ecosystem 4/5
  • Innovation 4/5
  • Support 3/5
  • Scalability 3/5
  • Trust 5/5
  • Demand (live) 1/5

The marker shows the average for the AI Coding & Dev category (20 tools)

Overall 3.7 / 5 · reviewed 2 Oct 2026

Pricing

Local use Free Free and open source - runs on your hardware
Cloud models Custom Paid tier - some models marked :cloud require an upgrade

Checked 15 Aug 2026. Pricing changes often; confirm on the official site before buying. Official pricing →

Verdict

Ollama is the least painful way to run a local model, and it works on ordinary hardware — we got 23 tokens per second on a CPU-only desktop, which is roughly reading speed. If you want to try local AI without buying a GPU, start here.

Who should not use it: anyone who needs instant answers. Our test prompt took 86 seconds. A hosted model answers the same question in two or three. If you are building something interactive, or you are impatient, local inference on a CPU will frustrate you.

Tested on

Date 15 August 2026
Ollama version 0.32.13
OS Windows 11 Pro 64-bit, build 26200
CPU Intel Core i5-12400 — 6 cores / 12 threads, 2.5 GHz base
RAM 32 GB
GPU Intel UHD Graphics 730 (integrated, 128 MB dedicated VRAM)
Model qwen3.5:2b — 2.7 GB download

There is no discrete graphics card in this machine. Inference ran on the CPU. That is the point of the test — almost every published benchmark assumes an NVIDIA card, and most people do not have one.

Measured performance

Generation speed 23.08 tokens/s
Prompt processing 83.77 tokens/s
Model load time 234 ms
Tokens generated 1,985
Total response time 1 min 26.5 s

The prompt was: “Explain what a vector database is, in three short paragraphs.”

The thing nobody warns you about

We asked for three short paragraphs. Ollama returned 1,985 tokens.

qwen3.5:2b is a reasoning model, and it thinks out loud. Before answering it produced an eight-step “Thinking Process”, drafted the paragraphs three times, second-guessed itself — “Wait, I need to make sure I don’t add unnecessary formatting”, “Actually, I can merge some ideas to make it punchier” — and ran a final check that it had in fact written three paragraphs.

The answer itself is around 260 tokens. Roughly 85% of the compute went to deliberation nobody asked for — about 75 seconds of visible thinking for 9 seconds of answer.

This is not a flaw in Ollama. It is a property of the model, and it is invisible until you run it. If you pick a reasoning model for a simple task, you pay for the reasoning. Choose accordingly.

What it actually costs

Free to install and run locally. No account required for local models.

Ollama’s launcher now also offers :cloud models. In our install, two of the three recommended models were cloud-hosted and one was marked “Upgrade required” — so there is a paid tier, and it is surfaced prominently in a tool whose premise is local execution. Worth knowing before you assume everything on the menu runs on your machine.

The real cost is hardware and disk. Models are large: a 2B model is ~2.7 GB, and the 26B Gemma 4 offered in the launcher is around 19 GB.

Pricing checked 15 August 2026.

Limits and failure modes

  • CPU-only inference is slow but usable. 23 tokens/s is about reading speed. Fine for drafting, painful for anything interactive.
  • Model names are not guessable. gemma4:4b does not exist — Gemma 4’s small variants are e2b and e4b. We also had llama3.2:3b fail. Check the tags before pulling.
  • Reasoning models are expensive on simple prompts — see above.
  • Disk fills quickly. Use ollama rm to clear models you have finished with.
  • No GPU acceleration on Intel integrated graphics in this configuration.

Alternatives

  • LM Studio — the same idea with a graphical interface. Better starting point if you dislike the terminal.
  • llama.cpp — the engine underneath much of this. More control, more work.
  • Open WebUI — a browser front-end that sits on top of Ollama.
  • Jan — offline-first desktop app, no terminal at all.

Install

Download from ollama.com/download and run the installer. Then:

ollama --version
ollama pull qwen3.5:2b
ollama run qwen3.5:2b --verbose

The --verbose flag prints the timing block after each response. Without it you get an answer and no data.

Full walkthrough, including the pulls that failed: Install Ollama and run your first local model

Last verified: 15 August 2026 · Ollama 0.32.13

Ready to try Ollama?

Visit the official website to get started — most tools have a free plan or free trial.

🚀 Try Ollama Now →
🔧
Need more free online tools? AMTake offers 150+ free tools — PDF tools, SEO tools, image compressors, converters & more. No signup required.
Explore Free →

AITechSpark

Premium AI-powered news covering AI, Digital Marketing, SaaS, Tech Tools, WordPress, SEO & Automation.

What is AITechSpark? →

Learn More