Using Codex CLI With Local or Other Models

Codex CLI can run on a model on your own machine through Ollama or LM Studio, or on another provider such as OpenRouter. The flags and config.toml entries that do it, why the provider must speak the Responses API, what gets worse, and when it is worth the trouble.

6 min read

Codex CLI runs on OpenAI’s models by default, but it can use a model on your own machine or from another provider. For a local model, start it with codex --oss and it talks to Ollama or LM Studio, both of which Codex knows by name. For anything else, such as OpenRouter or Mistral, you add a provider table to ~/.codex/config.toml with a base URL and the name of the environment variable that holds the key. There is one hard requirement: the provider must speak OpenAI’s Responses API, because that is the only protocol Codex supports today. Expect the agent to be less reliable with smaller models, to need a larger context window than local tools usually give by default, and to be slower on ordinary hardware.

What Codex supports

OpenAI’s advanced configuration page (opens in a new tab) describes three built-in provider IDs, openai, ollama and lmstudio, which are reserved, plus a built-in amazon-bedrock provider. Anything else is a custom provider that you define under [model_providers.<id>] and select with model_provider. The configuration reference lists the keys a provider takes: name, base_url, env_key, wire_api, query_params, http_headers, retry and timeout settings. For wire_api it lists responses as the only supported value and the default. A provider that only offers the older Chat Completions endpoint will not work, however “OpenAI-compatible” it says it is.

Ollama: the quick way and the manual way

Ollama can set Codex up for you. Ollama’s Codex page (opens in a new tab) describes ollama launch codex, which refreshes the model list and starts Codex with a dedicated profile; --config writes the setup without launching, and --restore removes it. The manual route is the --oss flag, with -m for the model:

Terminal: Codex CLI with Ollama
ollama pull gpt-oss:20b
codex --oss --local-provider ollama -m gpt-oss:20b

# or let Ollama write the profile and start Codex
ollama launch codex

The same Ollama page asks for a context window of at least 64k tokens for Codex. That is the setting most likely to catch you out, because Ollama sizes the default by video memory: under 24 GiB it gives a model 4k tokens, which is far too little for an agent that reads files and command output. Raise it in the Ollama app’s settings, or start the server with OLLAMA_CONTEXT_LENGTH=64000 ollama serve, then check the CONTEXT column of ollama ps while Codex is running.

LM Studio

LM Studio exposes the same Responses endpoint on port 1234. Start its server from the app or with lms server start --port 1234, then run codex --oss --local-provider lmstudio. The LM Studio guide (opens in a new tab) says Codex downloads and uses openai/gpt-oss-20b by default, that -m picks any other model you have loaded, and that the model should have more than about 25k tokens of context.

Make a local model the default, or keep it in a profile

Without a default, codex --oss asks which local provider to use, and codex exec --oss stops with an error instead of asking. Set oss_provider once. If you switch between OpenAI’s models and local ones, put the local settings in a profile: a separate file that Codex layers over config.toml when you pass --profile.

~/.codex/config.toml and ~/.codex/local.config.toml
# ~/.codex/config.toml
oss_provider = "ollama"            # or "lmstudio"

# ~/.codex/local.config.toml  (run: codex --profile local --oss)
model = "gpt-oss:20b"
model_context_window = 64000       # tell Codex what the server really allows
model_auto_compact_token_limit = 48000

The last two lines matter more locally than they do with OpenAI’s models. model_context_window tells Codex how much room it has, and model_auto_compact_token_limit makes it summarise the conversation before it runs out rather than after. Keep both at or below what Ollama or LM Studio actually allocates.

OpenRouter and other hosted providers

For a hosted provider, define it as a custom provider. OpenRouter documents (opens in a new tab) an OpenAI-compatible Responses API at https://openrouter.ai/api/v1/responses, so the base URL is the part before /responses. Keep the key in an environment variable and name it with env_key, never paste it into the file.

~/.codex/config.toml and ~/.codex/openrouter.config.toml
# ~/.codex/config.toml
[model_providers.openrouter]
name = "OpenRouter"
base_url = "https://openrouter.ai/api/v1"
env_key = "OPENROUTER_API_KEY"
wire_api = "responses"

# ~/.codex/openrouter.config.toml  (run: codex --profile openrouter)
model_provider = "openrouter"
model = "openai/gpt-oss-120b"

The same shape works for any provider with a Responses endpoint: OpenAI’s own example defines Mistral with base_url = "https://api.mistral.ai/v1" and env_key = "MISTRAL_API_KEY". OpenRouter notes that its Responses API is stateless, so every request carries the whole conversation; run a short task first to confirm the model you chose handles Codex’s tool calls. If all you need is to send OpenAI models through a proxy or a data-residency endpoint, you do not need a custom provider at all: set openai_base_url instead.

Check what Codex is actually using

  • /status at the start of a session shows the active model, the approval policy and token usage. If it names an OpenAI model when you expected a local one, the profile or flag did not apply.
  • /debug-config lists the configuration layers in order, so you can see whether a profile file, a project .codex/config.toml or your user config set the value that won.
  • A project’s .codex/config.toml cannot change provider settings: OpenAI’s documentation says project files may not redirect credentials or change provider auth. Keep provider tables in your user config or a profile.
  • If Codex waits for a long time and then fails, check the server first. ollama ps shows whether the model is loaded and how much of it runs on the GPU; a model that spills onto the CPU is much slower.

What gets worse

  • Tool calling. Codex works by calling tools in a loop: run a command, read a file, apply a patch, run the tests. A model that emits a malformed call, or forgets to call a tool at all, stalls the loop. Smaller local models tend to do this more often than the OpenAI models Codex’s documentation recommends.
  • Context. An agent reads a lot: instructions, files, command output. With 4k or 8k tokens it loses the task within a few steps. With 64k it copes; large repositories still need tight prompts.
  • Speed. Every step is a full model call. On a laptop GPU, a task that is quick with a hosted model can take much longer.
  • Features that belong to a ChatGPT sign-in. Cloud tasks, code review on GitHub and the Slack integration run on OpenAI’s side, and OpenAI’s pricing page lists none of them for API-key use. A local model changes what runs in your terminal, nothing else.
  • Web search. Through an Ollama profile, Codex’s web search requests are run by Ollama and need ollama signin. For custom providers, standalone web search is off unless the provider declares support.

When it is worth it

  • Code that must not leave the machine, for policy or contract reasons, and a machine with the memory to run a capable model.
  • Working offline, on a train or a locked-down network.
  • Routine, well-scoped jobs where a mid-sized model is enough: renaming, writing tests from a clear pattern, drafting documentation.
  • Comparing models from several vendors through one provider such as OpenRouter, without changing tools.
  • Not worth it: long, multi-file changes you need to be right first time. Use the models Codex is built around, and see Codex CLI best practices for keeping those runs tight.

A local model and a task board

MCP servers in config.toml work whatever the model provider, so a local Codex can still read and update a fenbs board at https://fenbs.ai/api/mcp. Whether it does so reliably depends on the same tool calling as everything else. Start with read-only tools listed in enabled_tools, let it report what is in Next Up, and add the tools that move tasks and set a test status once you trust its calls. Each change it makes is recorded in History under the assistant’s name, so a model that misfires is easy to spot. Connection steps are on Codex CLI on fenbs.

Related

Tool calling with local models in general: MCP with Ollama. The same setup for Anthropic’s agent: Claude Code with Ollama, LM Studio and OpenRouter. Every flag and config key: Codex CLI commands.

Questions people ask.

How do I use Codex CLI with Ollama?

Run ollama launch codex and Ollama sets Codex up for you, or run codex --oss --local-provider ollama -m followed by a model you have pulled. Give the model at least 64k tokens of context, because Ollama defaults to 4k on machines with less than 24 GiB of video memory.

Can Codex CLI use OpenRouter?

Yes. Define a custom provider in config.toml with base_url https://openrouter.ai/api/v1, env_key naming the variable that holds your OpenRouter key, and wire_api responses, then select it with model_provider. It works because OpenRouter offers a Responses API, which is the only protocol Codex supports.

Why does my OpenAI-compatible provider not work with Codex?

Codex supports only the Responses API. Its configuration reference lists responses as the only value for wire_api, so a provider that offers only Chat Completions will fail. Check that the provider documents a /responses endpoint.

Does a local model use my ChatGPT usage limits?

The model runs on your own machine, so its calls do not go to OpenAI. Cloud tasks, GitHub code review and Slack still need a ChatGPT sign-in and use your plan as usual.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.