MCP With Local Models via Ollama
Ollama runs the model and understands tool calls, but it does not connect to MCP servers itself. How people pair local models with MCP through a host, which hosts do it, where small models struggle with tools, and what stays private.
7 min read
Ollama is a model server, not an MCP host. It runs a model on your machine and, for models trained for it, returns tool calls through its API, but it does not connect to MCP servers or run their tools. To use MCP with a local model you put a host in between: an application that connects to MCP servers, hands their tools to the Ollama model as tool definitions, runs the calls the model asks for, and passes the results back. Open WebUI, Goose, Cline, and Claude Code or Codex pointed at Ollama all work this way. It is a good fit for private data and experiments, and a weaker one for long, many-tool agent work, where small models make more mistakes.
If MCP itself is new, MCP for beginners connects a first server in ten minutes. How MCP relates to a model’s own tool feature is in MCP vs function calling, and building an agent around tools in code is compared in MCP vs LangChain tools.
Who does what
- Ollama serves the model, by default on
127.0.0.1:11434. It accepts atoolslist with each chat request and replies withtool_callswhen the model wants one. - The host is the MCP client. It connects to each server, lists its tools, converts them into the tool definitions Ollama expects, and executes the calls.
- The MCP server does the work: reads files, queries a database, lists tasks on a board. It never talks to Ollama and cannot tell which model is on the other side.
Ollama’s tool-calling documentation (opens in a new tab) is explicit about the division: your application executes the tools, then sends each result back as a message with the role tool so the model can continue. It supports parallel calls and streaming, and describes the agent loop, where the model decides when to call a tool and folds the result into its reply. Every host below is, at heart, that loop plus an MCP client.
Hosts that pair Ollama with MCP
- Open WebUI. A browser chat interface that runs on your own server. Open WebUI’s MCP documentation (opens in a new tab) says native MCP support arrived in v0.6.31 and is Streamable HTTP only; you add a server under the admin settings as an external tool server, with no authentication, a bearer token or OAuth 2.1. Local stdio servers need the
mcpoproxy in front of them. - Goose. An open-source agent that treats its extensions as MCP servers and lists Ollama among its model providers. Its documentation says Goose relies heavily on tool calling, so a model without it can only chat and all extensions must be off.
- Cline and Codex. Both are MCP clients that run against Ollama models; Ollama’s web search page (opens in a new tab) shows the MCP configuration for Cline, Codex and Goose side by side. Codex starts on a local model with
codex --ossorollama launch codex. - Claude Code. Ollama’s Claude Code guide (opens in a new tab) runs it on an Ollama model with
ollama launch claude, or by pointingANTHROPIC_BASE_URLathttp://localhost:11434. Claude Code stays the MCP client, so servers you add withclaude mcp addare offered to the local model as tools.
ollama pull qwen3 OLLAMA_CONTEXT_LENGTH=64000 ollama serve # in another terminal: Claude Code on the local model ollama launch claude claude mcp add --transport http fenbs https://fenbs.ai/api/mcp
Whatever the host, pick a model built for tools. Ollama’s library marks models with tool support, and its own posts on tool calling name families such as Qwen 3, Llama 3.1 and Devstral. A model without tool support will answer in prose and never call anything.
Where small models struggle
Open WebUI’s documentation puts the central point in one line: MCP connects the application to your tools; it does not improve the model’s ability to use them. A connection that works with a large hosted model can fail with a small local one, and the failure is the model’s, not the protocol’s.
- Context runs out first. Ollama’s context length page (opens in a new tab) sets the default by memory: 4k tokens below 24 GiB of VRAM, 32k up to 48 GiB, 256k above. It recommends at least 64,000 tokens for agents, coding tools and web search. Every tool definition, and every tool result, sits in that window, and at 4k a handful of servers leaves little room for the conversation.
ollama psshows what a loaded model actually got. - The wrong tool, or none. Small models more often answer from memory, pick a tool with a similar name, or pass arguments that do not match the schema. Naming the tool in the prompt helps, as does a clear description on the server’s side.
- Long chains drift. Five dependent calls in a row, each reading the last result, is where a small model loses track. Break the job into steps you confirm.
- Memory and speed trade against each other. A bigger context uses more memory, and a model that spills from GPU to CPU slows sharply;
ollama psshows which processor it is running on.
The practical answer is to shrink the problem. Connect one or two servers, not ten. Where the host allows it, expose only the tools the job needs, such as Codex’s enabled_tools. Keep reads and writes in separate requests, and keep the host’s approval prompt on for anything that changes data.
Privacy: a local model is not a local workflow
Running the model yourself keeps your prompts and its answers on your machine. Ollama’s FAQ (opens in a new tab) says it does not see your prompts or data when you run locally, that it listens on 127.0.0.1 unless you change OLLAMA_HOST, and that OLLAMA_NO_CLOUD=1 turns off its cloud features. That matters, because models with a :cloud suffix run on Ollama’s servers rather than yours.
The MCP servers are a separate question. A remote server such as GitHub, Notion or a task board receives every argument the model sends it and returns your data from its own systems; that traffic is the same whether the model is local or hosted. What a local model changes is that the results, once fetched, are read on your machine rather than sent to a model provider. For a workflow where nothing leaves the machine, use local stdio servers, such as the reference Filesystem server limited to one folder, with a local model and cloud features off.
- Model local, servers local: the most private setup, and the one most limited by the model’s ability.
- Model local, servers remote: your prompts stay with you, but tool calls and their data go to each server’s owner under their terms.
- Model in the cloud, servers remote: the usual hosted setup, with a model provider in the path as well.
Where a task board fits
A board is a good first server for a local model because every write is visible. fenbs is a remote MCP server, so it sees the tool calls, not your prompts. A host that can open a browser signs in and gets a token that acts as you, narrowed by the scopes you tick; a host that cannot sends a token issued by hand under Settings, with a name, scopes and an optional expiry. Start with read only and a two-line prompt such as “call fenbs_whoami, then list the tasks in Next Up”. When the local model gets that right every time, tick write. Every change it makes is recorded in History under the assistant’s name, so a small model’s mistakes are easy to find and move back.
Related
Keeping any assistant on a short lead: MCP security best practices. Nine workflows to try once it works: MCP examples. Connecting a board: the MCP docs and what assistant tokens and scopes are.