Which version people actually mean
“Stable Diffusion” covers several generations with materially different characteristics, and conflating them causes most of the confusion around the name.
SD 1.5 remains the most widely used despite its age. It is small, runs on modest hardware, and has by far the deepest ecosystem of fine-tunes, LoRAs and ControlNet models. Its raw quality is dated; its ecosystem is not.
SDXL is larger, produces better composition and detail, and needs more VRAM. It has a substantial ecosystem of its own.
SD3 and later improved prompt adherence and text rendering, and arrived under different licence terms — which is the point most people miss and the reason it matters which one you download.
Practical consequence: a tutorial, a LoRA or a settings recommendation is written for a specific generation and often does not transfer. Check what a resource targets before following it.
Licensing, and why “open source” is too coarse a word
This is the section that decides whether you can use it for what you intend.
Earlier releases used permissive licences that placed few practical limits on commercial use. Later releases, including the SD3 generation, moved to a community licence with conditions — commercial use gated by revenue thresholds above which a separate agreement is required.
Three things follow, and all three catch people out.
Version matters more than family. “Stable Diffusion is open” is not a statement you can plan a business on; the terms of the exact model file are.
Fine-tunes inherit their base model’s terms. A community checkpoint built on a restrictively licensed base carries those restrictions, whatever the uploader wrote in the description.
The terms can change for future releases, though not retroactively for weights already published under earlier terms — which is a real argument for pinning a version you have verified rather than always taking the newest.
If the licensing burden is unworkable, FLUX schnell is Apache 2.0 and Adobe Firefly sells indemnification. Both are legitimate answers to a question Stable Diffusion cannot make simple.
The ecosystem, which is the actual product
Because the weights are downloadable, what has been built around them matters more than the base models themselves.
LoRAs teach a base model a specific style, character, object or concept from a small set of examples. They are typically tens of megabytes rather than gigabytes, load alongside a checkpoint, and can be combined. Tens of thousands exist publicly.
ControlNet conditions generation on a pose, depth map, edge map, scribble or segmentation. This is the single most important capability for anyone who needs a specific composition rather than a pleasing one — it converts a lottery into direction, and no closed hosted model offers an equivalent.
Community fine-tunes specialise the base model for illustration, photography, architecture, product rendering and much else, often outperforming the base model substantially within their niche.
None of this is possible with a model you can only call through an API, and it is the reason people accept the setup burden.
Training your own, and what it actually takes
Teaching the model your own subject — a product, a character, a house style — is the strongest reason to run local weights, and it is more accessible than people assume.
A LoRA can be trained from a few dozen consistent images on a consumer GPU in tens of minutes to a few hours. The quality of the result depends far more on the consistency of the source set than on any training parameter: varied lighting, varied crops and a clear subject beat a large messy collection every time.
If you lack a GPU, Leonardo AI offers hosted custom training that reaches a similar outcome without hardware.
Hardware, stated as numbers
- An interface — the weights alone do nothing.
- An NVIDIA GPU. Roughly: 6–8 GB VRAM runs SD 1.5 comfortably; SDXL wants 8–12 GB; newer generations more. AMD works on Linux with friction; Apple silicon works and is slower.
- Disk: 2–7 GB per checkpoint, and collections grow faster than expected.
- Enough understanding of prompting, samplers, steps and CFG to get past mediocre first results — which is a real learning curve, not a formality.
What it has no answer for
It is a set of weights. There is no interface, no support, no account and no guarantee. Quality varies wildly between community checkpoints and nothing ranks them for you.
Text inside images is weak on most versions — Ideogram exists largely because of that gap. Prompt adherence trails the leading hosted models on complex instructions. And training-data provenance is unresolved, which is a live consideration for commercial work and precisely what Firefly’s indemnified position is sold against.
Who should run it
Anyone needing control a hosted service will not give: a specific trained style, a reproducible pipeline, generation at volume without per-image cost, or work that cannot be uploaded to a third party.
Developers embedding generation into a product, where per-image API pricing does not survive contact with the business model.
People who want to train on their own material — the single strongest reason to be here at all.
It is a poor fit for someone who wants a good image quickly. A hosted service beats it on both quality and effort for casual use, and pretending otherwise wastes people’s evenings.
The case for local weights
- The largest ecosystem in image generation — LoRAs, fine-tunes and ControlNet in the tens of thousands.
- Trainability on your own subject, which no closed model allows.
- No per-image cost and no uploads, which matters for volume and for confidential work.
- Reproducibility — the same seed, model and settings give the same image indefinitely.
- No content policy but your own, and no vendor able to change the model under you.
- Version pinning, so a working setup keeps working.
What you take on
- Licence terms differ by version — the most common and most expensive misunderstanding.
- Base quality trails leading hosted models, particularly on prompt adherence.
- Weak text rendering on most versions.
- You are the support. No help desk, and community checkpoint quality is unranked.
- Training-data provenance is unresolved, which matters commercially.
- A real learning curve before results justify the effort.
Nothing to leave
There is no vendor to leave. Weights are files on your disk, LoRAs are files, and every interface in this family reads the same formats — switch from A1111 to ComfyUI and your models come with you unchanged.
That permanence is a genuine and underrated property. A hosted model can be deprecated, retuned or withdrawn; a checkpoint on your disk produces the same image in five years. For anyone whose work depends on reproducibility, that alone justifies the setup cost.
The alternatives
- FLUX — the current open-weights alternative, better prompt adherence, clearer per-variant licensing.
- Midjourney — better images with no control and no local option.
- Adobe Firefly — licensed training data with commercial indemnification.
- Leonardo AI — hosted access to custom training without running hardware.
Compiled from Stability AI’s documentation and public sources. Licence terms differ by model version — verify against the specific release you intend to use. We have not hands-on tested this tool. Last reviewed 16 August 2026.