Using Codex CLI With Local or Other Models
Codex CLI can run on a model on your own machine through Ollama or LM Studio, or on another provider such as OpenRouter. The flags and config.toml entries that do it, why the provider must speak the Responses API, what gets worse, and when it is worth the trouble.
6 min read
Codex CLI runs on OpenAI’s models by default, but it can use a model on your own machine or from another provider. For a local model, start it with codex --oss and it talks to Ollama or LM Studio, both of which Codex knows by name. For anything else, such as OpenRouter or Mistral, you add a provider table to ~/.codex/config.toml with a base URL and the name of the environment variable that holds the key. There is one hard requirement: the provider must speak OpenAI’s Responses API, because that is the only protocol Codex supports today. Expect the agent to be less reliable with smaller models, to need a larger context window than local tools usually give by default, and to be slower on ordinary hardware.
What Codex supports
OpenAI’s advanced configuration page (opens in a new tab) describes three built-in provider IDs, openai, ollama and lmstudio, which are reserved, plus a built-in amazon-bedrock provider. Anything else is a custom provider that you define under [model_providers.<id>] and select with model_provider. The configuration reference lists the keys a provider takes: name, base_url, env_key, wire_api, query_params, http_headers, retry and timeout settings. For wire_api it lists responses as the only supported value and the default. A provider that only offers the older Chat Completions endpoint will not work, however “OpenAI-compatible” it says it is.
Ollama: the quick way and the manual way
Ollama can set Codex up for you. Ollama’s Codex page (opens in a new tab) describes ollama launch codex, which refreshes the model list and starts Codex with a dedicated profile; --config writes the setup without launching, and --restore removes it. The manual route is the --oss flag, with -m for the model:
ollama pull gpt-oss:20b codex --oss --local-provider ollama -m gpt-oss:20b # or let Ollama write the profile and start Codex ollama launch codex
The same Ollama page asks for a context window of at least 64k tokens for Codex. That is the setting most likely to catch you out, because Ollama sizes the default by video memory: under 24 GiB it gives a model 4k tokens, which is far too little for an agent that reads files and command output. Raise it in the Ollama app’s settings, or start the server with OLLAMA_CONTEXT_LENGTH=64000 ollama serve, then check the CONTEXT column of ollama ps while Codex is running.
LM Studio
LM Studio exposes the same Responses endpoint on port 1234. Start its server from the app or with lms server start --port 1234, then run codex --oss --local-provider lmstudio. The LM Studio guide (opens in a new tab) says Codex downloads and uses openai/gpt-oss-20b by default, that -m picks any other model you have loaded, and that the model should have more than about 25k tokens of context.
Make a local model the default, or keep it in a profile
Without a default, codex --oss asks which local provider to use, and codex exec --oss stops with an error instead of asking. Set oss_provider once. If you switch between OpenAI’s models and local ones, put the local settings in a profile: a separate file that Codex layers over config.toml when you pass --profile.
# ~/.codex/config.toml oss_provider = "ollama" # or "lmstudio" # ~/.codex/local.config.toml (run: codex --profile local --oss) model = "gpt-oss:20b" model_context_window = 64000 # tell Codex what the server really allows model_auto_compact_token_limit = 48000
The last two lines matter more locally than they do with OpenAI’s models. model_context_window tells Codex how much room it has, and model_auto_compact_token_limit makes it summarise the conversation before it runs out rather than after. Keep both at or below what Ollama or LM Studio actually allocates.
OpenRouter and other hosted providers
For a hosted provider, define it as a custom provider. OpenRouter documents (opens in a new tab) an OpenAI-compatible Responses API at https://openrouter.ai/api/v1/responses, so the base URL is the part before /responses. Keep the key in an environment variable and name it with env_key, never paste it into the file.
# ~/.codex/config.toml [model_providers.openrouter] name = "OpenRouter" base_url = "https://openrouter.ai/api/v1" env_key = "OPENROUTER_API_KEY" wire_api = "responses" # ~/.codex/openrouter.config.toml (run: codex --profile openrouter) model_provider = "openrouter" model = "openai/gpt-oss-120b"
The same shape works for any provider with a Responses endpoint: OpenAI’s own example defines Mistral with base_url = "https://api.mistral.ai/v1" and env_key = "MISTRAL_API_KEY". OpenRouter notes that its Responses API is stateless, so every request carries the whole conversation; run a short task first to confirm the model you chose handles Codex’s tool calls. If all you need is to send OpenAI models through a proxy or a data-residency endpoint, you do not need a custom provider at all: set openai_base_url instead.
Check what Codex is actually using
/statusat the start of a session shows the active model, the approval policy and token usage. If it names an OpenAI model when you expected a local one, the profile or flag did not apply./debug-configlists the configuration layers in order, so you can see whether a profile file, a project.codex/config.tomlor your user config set the value that won.- A project’s
.codex/config.tomlcannot change provider settings: OpenAI’s documentation says project files may not redirect credentials or change provider auth. Keep provider tables in your user config or a profile. - If Codex waits for a long time and then fails, check the server first.
ollama psshows whether the model is loaded and how much of it runs on the GPU; a model that spills onto the CPU is much slower.
What gets worse
- Tool calling. Codex works by calling tools in a loop: run a command, read a file, apply a patch, run the tests. A model that emits a malformed call, or forgets to call a tool at all, stalls the loop. Smaller local models tend to do this more often than the OpenAI models Codex’s documentation recommends.
- Context. An agent reads a lot: instructions, files, command output. With 4k or 8k tokens it loses the task within a few steps. With 64k it copes; large repositories still need tight prompts.
- Speed. Every step is a full model call. On a laptop GPU, a task that is quick with a hosted model can take much longer.
- Features that belong to a ChatGPT sign-in. Cloud tasks, code review on GitHub and the Slack integration run on OpenAI’s side, and OpenAI’s pricing page lists none of them for API-key use. A local model changes what runs in your terminal, nothing else.
- Web search. Through an Ollama profile, Codex’s web search requests are run by Ollama and need
ollama signin. For custom providers, standalone web search is off unless the provider declares support.
When it is worth it
- Code that must not leave the machine, for policy or contract reasons, and a machine with the memory to run a capable model.
- Working offline, on a train or a locked-down network.
- Routine, well-scoped jobs where a mid-sized model is enough: renaming, writing tests from a clear pattern, drafting documentation.
- Comparing models from several vendors through one provider such as OpenRouter, without changing tools.
- Not worth it: long, multi-file changes you need to be right first time. Use the models Codex is built around, and see Codex CLI best practices for keeping those runs tight.
A local model and a task board
MCP servers in config.toml work whatever the model provider, so a local Codex can still read and update a fenbs board at https://fenbs.ai/api/mcp. Whether it does so reliably depends on the same tool calling as everything else. Start with read-only tools listed in enabled_tools, let it report what is in Next Up, and add the tools that move tasks and set a test status once you trust its calls. Each change it makes is recorded in History under the assistant’s name, so a model that misfires is easy to spot. Connection steps are on Codex CLI on fenbs.
Related
Tool calling with local models in general: MCP with Ollama. The same setup for Anthropic’s agent: Claude Code with Ollama, LM Studio and OpenRouter. Every flag and config key: Codex CLI commands.