Open-Source LLMs: What Teams Use and Why
Most models called open source are open-weight: you can download and run them, but not rebuild them. What the Open Source AI Definition requires, the model families as their own makers describe them, how their licenses differ, and how to run one on your own machine.
7 min read
An open source LLM, in everyday use, is a language model whose trained weights you can download and run on your own hardware, under a license that lets you use and usually change it. Teams pick one to keep data on their own machines, to work offline, to fine-tune a model on their own material, or to stop depending on one vendor’s API. The best known families are Meta’s Llama, Google’s Gemma, Mistral’s open models, Alibaba’s Qwen, DeepSeek, OpenAI’s gpt-oss, Microsoft’s Phi and IBM’s Granite. Strictly, most of them are open-weight rather than open source, and the license is what decides what you may do with each.
What a language model is, in plain words, is in what is an LLM. This page is about the ones you can hold.
Open source AI vs open weights
The Open Source Initiative, which maintains the definition of open source software, published the Open Source AI Definition (opens in a new tab), now at version 1.0. It asks that an AI system be available under terms that grant four freedoms: to use it for any purpose, to study how it works, to modify it, and to share it with or without changes. To exercise those freedoms you need what it calls the preferred form to make modifications, which has three parts.
- Data information: enough detail about the training data that a skilled person could build a substantially equivalent system, including where the data came from and how it was selected and filtered.
- Code: the complete source code used to train and run the system.
- Parameters: the model weights and other settings, under terms the OSI approves.
The definition is direct about the gap: “Open Source models” and “Open Source weights” must include the data information and code used to derive those parameters. A downloadable file of weights, however permissive its license, does not meet that bar on its own. That is why careful writers say open-weight for most of the models below, and why you will see vendors use the words differently. Microsoft, for example, says its Phi models “are open source through the MIT License.” Both statements can be true in their own terms; the OSI test is stricter.
The open source AI models, as their makers describe them
No rankings here: which one is best depends on your hardware, your language and your task, and the leaders change month to month. These are the families and what each vendor’s own pages say, as of October 1, 2026.
- Meta Llama. llama.com now redirects to Meta’s developer site, which leads with Meta’s newer Muse models. Its Llama section lists Llama 4, “natively multimodal models with a mixture-of-experts architecture,” in Scout and Maverick versions, alongside Llama 3.x, available directly from Meta or through Hugging Face or Kaggle. They ship under the Llama 4 Community License (opens in a new tab), not a standard open source license.
- Google Gemma. Google calls Gemma “a family of lightweight, state-of-the-art open models” built from the same research as Gemini. The Gemma docs (opens in a new tab) list Gemma 4 as the current generation and link its license, Apache 2.0; earlier Gemma models remain under Google’s Gemma Terms of Use.
- Mistral. Mistral says it “develops, or makes available, open-weight and commercial large language models.” Its models page marks Mistral Large 3, Mistral Small 4 and the Ministral 3 sizes Apache 2.0, Mistral Medium 3.5 as open weights under a Modified MIT license, and others, such as Codestral, as Premier.
- Alibaba Qwen. The Qwen team’s Qwen3.8 repository, under Apache 2.0, introduces Qwen3.8-27B and a large mixture-of-experts model, with weights on Hugging Face and ModelScope. Other Qwen models carry other licenses, so read each model card.
- DeepSeek. DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026, with weights on Hugging Face under the MIT license, and serves the same family through its API.
- OpenAI gpt-oss. OpenAI calls gpt-oss-120b (opens in a new tab) its most powerful open-weight model, one that fits on a single H100 GPU, and gpt-oss-20b a medium-sized open-weight model for low latency. Both are Apache 2.0.
- Microsoft Phi. Microsoft describes Phi as its family of small language models, made so developers can run AI on a device without a cloud connection; Phi-4-mini and Phi-4-multimodal are the newest, under the MIT license.
- IBM Granite. IBM presents Granite 4.2 as lightweight models “released under an Apache 2.0 license,” cryptographically signed, with an ISO certification for the management system behind the language models.
Licenses in plain words
The family name tells you little. The license on the exact model you download tells you what you may do. They fall into four groups.
- Permissive licenses: Apache 2.0 and MIT. Use, change and ship the model commercially, keep the notices, and you are largely done. Apache 2.0 also includes an explicit patent grant. gpt-oss, Gemma 4, Granite, Phi, DeepSeek-V4.1-Flash, Qwen3.8 and several Mistral models fall here.
- Modified permissive licenses, such as Mistral’s Modified MIT. Read the modification; it is the part that matters.
- Community licenses written by the vendor, such as Llama’s. The Llama 4 license asks companies with more than 700 million monthly active users to request a separate license, requires “Built with Llama” to be displayed, binds you to an acceptable use policy, and requires a model you train from Llama materials and distribute to have “Llama” at the start of its name.
- Non-commercial licenses. Some models in otherwise open lineups are research or non-commercial only; Mistral lists Voxtral TTS under CC BY-NC 4.0. Fine for experiments, not for a product.
Two habits save trouble. Record the exact model name, version and license when you adopt one, because licenses change between generations, as Gemma’s did. And read the acceptable use or prohibited use policy; most vendors have one even when the license is permissive.
Running an open source LLM locally
Three tools cover most local setups. Ollama downloads and runs models with one command and serves them over a local API. LM Studio is a desktop app for Apple Silicon Macs, x64 and ARM64 Windows PCs and x64 Linux, running GGUF models through llama.cpp and MLX models on Apple Silicon, with a local server other tools can call. And llama.cpp (opens in a new tab), the C/C++ engine under much of this, runs a model straight from Hugging Face and starts an OpenAI-compatible server.
# Ollama: download and chat with a small Gemma 4 model ollama run gemma4:e2b # llama.cpp: run a model from Hugging Face, or serve it llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF # OpenAI-compatible API
Start small. A model of a few billion parameters runs on a recent laptop; a mixture-of-experts model of hundreds of billions needs server GPUs, however few parameters are active per token. Check the model card for memory needs before downloading tens of gigabytes. Pointing a coding agent at a local model is covered in Claude Code with Ollama, and giving a local model tools through a host is in MCP with Ollama.
When to choose an open source LLM
- Data must stay on your hardware, for a client contract, a regulation or your own policy.
- You need to work offline, or on a network that cannot reach a cloud API.
- You want to fine-tune on your own material and keep the result.
- The job is narrow and repeated, such as classifying tickets or extracting fields, where a small model does well and the volume would make API calls add up.
- You need the model to stay the same for years, rather than change when a vendor retires a version.
And when not to: long, many-step agent work and hard reasoning still favor the largest hosted models, and running your own means you patch, monitor and secure it. Many teams run both, a local model for private, bounded jobs and a hosted one for the rest.
A checklist before you adopt one
- Name the job in one sentence and test two or three candidates on twenty real examples of it.
- Write down the exact model, version and license, and who checked the license.
- Confirm the hardware: memory for the weights plus room for the context you need.
- Check whether the model supports tool calling if an agent will use it.
- Decide who updates it, and how you will know a newer version exists.
- Keep a hosted model as a fallback for the jobs the local one cannot do.
Recording the choice
Which model you run, under which license, is a decision someone on the team made, and the next person or assistant needs to find it. In fenbs that goes on the Decisions and rules page: the person who decided, the reason, and, if it should hold from now on, marked as a rule such as “customer data is only processed by the local model.” Every assistant connected to the board over MCP reads the rules first, through fenbs_get_context, before it starts work. fenbs does not host or run models and cannot tell which one an assistant used; it keeps the decision and the tasks, and records in History who changed what.
Related
Building on a model with a framework: what is LangChain. Tools that write code with these models: AI code generators. Connecting an assistant to a board: the MCP docs and AI context.