Claude Code With Local Models: Ollama, LM Studio and OpenRouter

Claude Code can be pointed at a model other than Claude by changing one environment variable. How Ollama, LM Studio and OpenRouter document it, what Anthropic does and does not support, what breaks, and what stays private.

7 min read

Claude Code speaks one protocol, the Anthropic Messages API, and sends it to whatever address ANTHROPIC_BASE_URL names. Ollama and LM Studio now serve an Anthropic-compatible endpoint on your own machine, and OpenRouter serves one in the cloud, so each can sit behind Claude Code with three environment variables and a --model flag. It works, and each vendor documents it. Anthropic documents the variable and the gateway pattern, but says plainly that it does not support routing Claude Code to non-Claude models. So the setup is yours to own: expect weaker tool use, a tight context window and a few features that switch off, and use a local model for the jobs where those costs are worth paying.

This is about running Claude Code itself on another model. Pairing MCP servers with a local model in general, through hosts such as Open WebUI and Goose, is covered in MCP with local models via Ollama.

What Anthropic documents, and what it does not

Anthropic’s guide to connecting Claude Code to an LLM gateway (opens in a new tab) is the reference for the moving parts. ANTHROPIC_BASE_URL points Claude Code at the gateway. The credential goes in ANTHROPIC_AUTH_TOKEN when the gateway wants a bearer token in the Authorization header, or in ANTHROPIC_API_KEY when it wants x-api-key. You can export both in your shell or put them in the env block of ~/.claude/settings.json; never in a project’s .claude/settings.json, which is committed. Run /status inside Claude Code and the Status tab shows the base URL and which credential is active.

The same documentation draws the line. Its LLM gateway page (opens in a new tab) says Anthropic does not endorse, maintain or audit third-party gateway products, and does not support routing Claude Code to non-Claude models through any gateway. A gateway in front of Claude, for cost control or audit logging, is a supported pattern. A local open-weight model behind the same variable is something the model vendors support, not Anthropic, and every setup below should be read that way.

Ollama

Ollama’s Claude Code guide (opens in a new tab) offers a one-line route, ollama launch claude, which picks a model and starts Claude Code against it. The manual route is three variables and a model name. Ollama asks for a context length of 64k or more for larger repositories, and its library tags which models support tools.

Terminal: Claude Code on a local Ollama model
# the quick route
ollama launch claude

# the manual route
export ANTHROPIC_BASE_URL=http://localhost:11434
export ANTHROPIC_AUTH_TOKEN=ollama     # required by Claude Code, ignored by Ollama
export ANTHROPIC_API_KEY=""
claude --model qwen3.5

Ollama implements a subset of the API, and its Anthropic compatibility page lists what is missing: tool_choice, prompt caching, the token-counting endpoint and PDF documents are not supported, and deferred tools are not fully supported. Messages, streaming and tool calls work with compatible models. Models with a :cloud suffix run on Ollama’s servers, not your machine, which matters for the privacy section below.

LM Studio

LM Studio exposes an Anthropic-compatible /v1/messages endpoint from its local server, by default on port 1234. LM Studio’s Claude Code page (opens in a new tab) sets the same two variables plus CLAUDE_CODE_ATTRIBUTION_HEADER=0, which in Claude Code’s own variable list omits the attribution block, carrying the client version and a prompt fingerprint, from the start of the system prompt. It recommends a model loaded with more than about 25k of context, because Claude Code uses a lot of it. If you switched on Require Authentication in LM Studio, the token you create there replaces the placeholder.

Terminal: Claude Code on LM Studio
lms server start --port 1234

export ANTHROPIC_BASE_URL=http://localhost:1234
export ANTHROPIC_AUTH_TOKEN=lmstudio
export CLAUDE_CODE_ATTRIBUTION_HEADER=0
claude --model openai/gpt-oss-20b

OpenRouter

OpenRouter is not local. It is a hosted gateway with an Anthropic-compatible endpoint that routes to many providers, which is why people reach for it to try other models in Claude Code. OpenRouter’s Claude Code guide (opens in a new tab) sets ANTHROPIC_BASE_URL to https://openrouter.ai/api, the OpenRouter key in ANTHROPIC_AUTH_TOKEN, and ANTHROPIC_API_KEY explicitly empty, then asks you to run /logout once if you were signed in with a Claude account, so the cached login does not conflict. Its own warning is worth repeating: Claude Code through OpenRouter is only guaranteed to work with the Anthropic first-party provider.

To choose models, override the class variables Claude Code resolves its aliases through, such as ANTHROPIC_DEFAULT_SONNET_MODEL and CLAUDE_CODE_SUBAGENT_MODEL, or set CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1 to fill the /model picker from the gateway. Billing is with OpenRouter, not your Claude plan.

What breaks

  • Tool calling is the whole job. Reading files, editing them and running commands are all tool calls, so a model without tool support can only chat. Smaller models that do support tools still pick the wrong one, pass arguments that do not match the schema, or stop halfway through a chain of edits.
  • Context fills fast. Claude Code’s instructions and tool definitions take a large share of a small window before you type anything. With ANTHROPIC_BASE_URL on a non-Anthropic host, MCP tool search is off by default, so every MCP tool definition loads up front as well. Connect fewer servers, and set CLAUDE_CODE_MAX_CONTEXT_TOKENS to the model’s real window so Claude Code compacts before the server truncates.
  • Some features need a claude.ai sign-in and Anthropic’s API. Remote Control is disabled while the base URL points at another host, and cloud sessions always run on Anthropic’s API, so neither works through Ollama, LM Studio or OpenRouter. See Claude Code remote control for what those features need.
  • Speed and cost behave differently. Without prompt caching, every turn resends the full context, and a local model that spills from GPU to CPU slows sharply. Long sessions feel it most.
  • Quality is the model’s, not the harness’s. Claude Code’s prompts and habits were written for Claude. Other models follow them less reliably, so plan mode, subagents and long autonomous runs degrade first.

When a local model is worth it

A local model earns its place on short, well-bounded jobs where privacy or being offline matters more than raw ability: explaining an unfamiliar file, writing docstrings, drafting tests for one function, renaming across a small module, or working on a plane. It is a poor fit for multi-file refactors, debugging that needs many dependent steps, and anything you would leave running unattended. Most people who keep one end up with two profiles and pick per task.

~/.bashrc or ~/.zshrc: a local profile beside the default
# normal sessions use your Claude sign-in; this one uses Ollama
claude-local() {
  ANTHROPIC_BASE_URL=http://localhost:11434 \
  ANTHROPIC_AUTH_TOKEN=ollama \
  ANTHROPIC_API_KEY="" \
  CLAUDE_CODE_MAX_CONTEXT_TOKENS=64000 \
  claude --model qwen3.5 "$@"
}

Setting the variables inline for one command, rather than in ~/.claude/settings.json, keeps the gateway from applying to every session and every background agent. If you do put them in settings, the settings value wins over the shell, so check /status when a session seems to be on the wrong model.

Privacy: what still leaves the machine

With Ollama or LM Studio on your own hardware, your prompts, code and the model’s replies stay local. Claude Code itself still sends operational traffic. Anthropic’s data usage page (opens in a new tab) describes metrics, which it says never include your code, prompts or file paths, and error reports with known secrets and paths redacted. CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 turns that traffic off, and also turns off auto-updates, so plan another way to update. The WebFetch tool’s domain safety check still calls Anthropic unless you set skipWebFetchPreflight in settings.

  • Ollama :cloud models and OpenRouter send your prompts and code to someone else’s servers, under their terms. That is not a local workflow, whatever the terminal looks like.
  • Remote MCP servers receive every tool call the model makes and return data from their own systems, whichever model is on the other end.
  • A local server on localhost should stay there. Ollama’s local server does not check API keys, so do not open it to your network to share a model with a colleague; LM Studio has a Require Authentication switch if you must.

Keeping the work visible with a weaker model

A smaller model makes more mistakes, so it pays to have its work land somewhere you can check. A fenbs board is a remote MCP server, and Claude Code stays the MCP client whichever model it runs on: claude mcp add --transport http fenbs https://fenbs.ai/api/mcp, then /mcp to sign in, and the board’s tools are offered to the local model like any other. Tick read and comment only at first, so it can report on tasks without moving them. Every change it makes is recorded in the board’s History as your assistant acting for you, so a wrong move is easy to spot and put back. The CLAUDE.md lines that make each session read the board first are in a task-tracking workflow for Claude Code.

Related

Keeping token use down on any model: reduce Claude Code token usage. Rolling a gateway out to a whole team: Claude Code for teams. Connecting a board: the Claude Code integration and assistant tokens and scopes.

Questions people ask.

Can Claude Code run on a local model with Ollama?

Yes, as Ollama documents it. Run ollama launch claude, or set ANTHROPIC_BASE_URL to http://localhost:11434, ANTHROPIC_AUTH_TOKEN to ollama and ANTHROPIC_API_KEY to empty, then start claude with --model and a tool-capable model. Anthropic does not support routing Claude Code to non-Claude models, so quality and compatibility are down to the model.

Which local models work with Claude Code?

Models trained for tool calling, loaded with a large context. Ollama recommends 64k tokens or more for larger repositories and LM Studio more than about 25k. A model without tool support cannot read or edit files through Claude Code at all.

Does Claude Code with OpenRouter use my Claude subscription?

No. With the OpenRouter key in ANTHROPIC_AUTH_TOKEN, requests carry that key and are billed by OpenRouter. OpenRouter recommends running /logout once so a cached Claude login does not conflict, and says the setup is only guaranteed with the Anthropic first-party provider.

Is Claude Code with a local model fully private?

Prompts and code stay on your machine, but Claude Code still sends operational metrics and error reports to Anthropic unless you set CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1, and remote MCP servers still receive every tool call. Cloud-suffixed Ollama models and OpenRouter run on other companies’ servers.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.