Codex Subagents and Multi-Agent Runs

Codex can split a job across subagents inside one session, run several cloud chats side by side, and be driven in parallel from a script. What each option does, how to set it up, and which one fits the work in front of you.

8 min read

Yes. Codex CLI can spawn subagents, and current releases have the feature switched on. OpenAI’s Codex subagents documentation (opens in a new tab) says subagent workflows are enabled by default in the CLI, the IDE extension and the ChatGPT desktop app. There is one catch people miss: at most intelligence levels Codex will not split the work on its own. You ask for it (“spawn two agents”, “do this in parallel”), or an AGENTS.md or skill tells it to. Each subagent works in its own thread and hands a summary back to the main one. For work that outgrows one session there are three more routes: parallel cloud chats, several Codex runs in separate Git worktrees, and the Codex SDK.

This page is about Codex. Running agents from different vendors side by side, and keeping them off each other’s branches, is covered in orchestrating coding agents. If you are comparing with Claude Code’s version, see Claude Code subagent examples.

What a subagent is in Codex

A subagent is a second agent thread that the main session starts for a piece of the job. It reads, searches and runs tools in its own context, and only its summary returns to your chat. OpenAI’s reason for the design is context: exploration output, test logs and dead ends stay in the subagent’s thread instead of filling the main one.

  • Codex ships three built-in agents: default for general work, worker for implementation, and explorer for read-heavy exploration of a codebase.
  • Subagents inherit the parent’s sandbox policy and permission mode. A change you make with /permissions during the session is applied to the agents it spawns afterwards.
  • Approval requests can come from a subagent thread you are not looking at. In the CLI they surface while you are on the main thread, so read which thread is asking before you approve.
  • A subagent inherits the parent’s model and reasoning effort unless its own configuration says otherwise.
  • They cost more. The documentation says plainly that because each subagent does its own model and tool work, a subagent workflow uses more tokens than a comparable single-agent run.

The one exception to “only when asked” is the Ultra intelligence level, which the documentation says uses maximum reasoning and lets Codex delegate suitable work to subagents without being told.

Asking for subagents

Name the split, say what each agent should look for, and say what you want back. This is close to the example OpenAI gives:

A prompt in a Codex CLI session
Review this branch with parallel subagents. Spawn one subagent for
security risks, one for test gaps, and one for maintainability. Wait for
all three, then summarise the findings by category with file references.

The strongest fit is work that splits cleanly and mostly reads: reviewing, tracing a bug through several modules, surveying how a library is used. When two pieces would edit the same files, it is simpler to run them one after the other, or to give each its own worktree as described below. And if you want delegation every time for a kind of task, write the instruction into AGENTS.md (“for reviews, spawn one explorer per package”) rather than repeating it in prompts.

Custom agents in TOML

Your own agents are standalone TOML files, one agent per file. Personal ones go in ~/.codex/agents/ and project ones in .codex/agents/, where they can be committed with the code. Three fields are required: name, description and developer_instructions. Optional fields include model, model_reasoning_effort, sandbox_mode, mcp_servers and skills.config.

.codex/agents/security-reviewer.toml
name = "security_reviewer"
description = "Read-only reviewer that looks for security risks in a diff."
sandbox_mode = "read-only"
model_reasoning_effort = "high"
developer_instructions = """
Review only. Do not edit files or run commands that change state.
Check input validation, auth checks, secrets in code and unsafe shell calls.
Report each finding with the file, the line and a one-line fix.
"""

Leaving model out lets the agent inherit whatever the session uses, which ages better than pinning a model name. Setting sandbox_mode = "read-only" on a reviewer is the useful habit here: the parent may be allowed to write, but the agent whose only job is to read cannot. Ask for it by name, as in “spawn the security_reviewer agent on the changes in src/auth”.

The [agents] settings

Session-wide behaviour lives in an [agents] table in ~/.codex/config.toml. The documented keys are few:

~/.codex/config.toml
[agents]
enabled = true                              # multi-agent tools on (the default)
max_concurrent_threads_per_session = 3      # cap on spawned threads at once
default_subagent_reasoning_effort = "medium"
interrupt_message = true                    # record an interruption message (the default)
  • enabled = false turns the multi-agent tools off, which is worth doing for a CI profile where you want one predictable thread.
  • max_concurrent_threads_per_session caps how many spawned threads run at once. Leave it unset and Codex picks the default; the documentation does not give a number.
  • default_subagent_model and default_subagent_reasoning_effort set what spawned agents use when neither the prompt nor the agent file says.
  • max_threads is a legacy alias for the concurrency cap. Rename it when you next touch the file.

While agents are running, /agent in the CLI switches between their threads so you can inspect what each one is doing.

Parallel cloud chats

Subagents run on your machine inside one session. When the jobs are long and independent, Codex cloud is often the better split. OpenAI now calls cloud tasks “cloud chats”: each runs in its own cloud environment, several can run at once, and you review each diff when it is ready. The Codex commands reference (opens in a new tab) lists the terminal side:

Terminal
codex cloud exec --env <ENV_ID> "Add pagination to the orders endpoint"
codex cloud exec --env <ENV_ID> --attempts 3 "Fix the flaky upload test"   # best-of-N, 1 to 4
codex cloud list                 # recent cloud chats (--json for scripts)
codex apply <TASK_ID>            # apply a finished chat's diff locally

--attempts is a different kind of parallelism from subagents: the same task tried up to four times, so you can keep the best result. Environments, internet access and review are covered in Codex cloud.

The desktop app has a local version of the same idea. Its worktrees page (opens in a new tab) says worktrees let Codex run several independent chats in one project without interfering, and it keeps them under $CODEX_HOME/worktrees. That page describes the app; the next section does the same by hand from the terminal.

Several Codex runs in worktrees

The plainest multi-agent setup needs no special feature at all: one Git worktree per task and one codex exec in each. Every run gets its own branch and files, so nothing collides, and each writes its last message to a file you can read afterwards.

Terminal (bash)
git worktree add ../app-auth -b agent/auth-timeout
git worktree add ../app-docs -b agent/docs-export

(cd ../app-auth && codex exec --sandbox workspace-write \
  -o ../auth-result.txt "Fix the session timeout bug in BUG-014") &
(cd ../app-docs && codex exec --sandbox workspace-write \
  -o ../docs-result.txt "Update the API docs for the export endpoint") &
wait

Add --json when a script needs the event stream rather than the final message, and --ephemeral when you do not want the run saved as a session. Review each branch as you would a colleague’s, then merge in an order you choose. Habits for keeping those runs reviewable are in Codex CLI best practices.

Driving Codex from code

When the orchestration needs logic of its own, such as retries, a queue, or results fed into the next step, the Codex SDK (opens in a new tab) controls local Codex agents from TypeScript (@openai/codex-sdk) or Python (openai-codex). A script can start one thread per worktree and wait for all of them:

run.mts (TypeScript)
import { Codex } from "@openai/codex-sdk";

const codex = new Codex();
const jobs = [
  { dir: "../app-auth", prompt: "Fix the session timeout bug in BUG-014" },
  { dir: "../app-docs", prompt: "Update the API docs for the export endpoint" },
];

const results = await Promise.all(
  jobs.map((job) =>
    codex.startThread({ workingDirectory: job.dir, sandboxMode: "workspace-write" }).run(job.prompt),
  ),
);
results.forEach((turn, i) => console.log(jobs[i].dir, turn.finalResponse));

One older pattern no longer works. Tutorials that run Codex as an MCP server and drive it from the OpenAI Agents SDK depend on codex mcp-server, which OpenAI has removed; integrations now go through the Codex app server, which speaks its own JSON-RPC protocol. The details are in MCP with OpenAI. Codex is still an MCP client, so it can connect to other servers.

Which one to use

  • Subagents: one job with independent, mostly read-only parts, where you want a single summary back in the chat you are in.
  • Cloud chats: long, separate jobs you want off your machine, or the same job attempted several times.
  • Worktrees with codex exec: separate jobs that all edit code locally, each on its own branch.
  • The Codex SDK: when a program, not a person, decides what runs next.

One list of work across every thread

Subagents report to the session that spawned them, and cloud chats and worktree runs report to nobody but you. None of them knows what the others took. A board both people and assistants can reach is the simplest fix. Add fenbs to Codex with codex mcp add fenbs --url https://fenbs.ai/api/mcp, and each run can read its task, move it to In Progress, comment what it changed and move it to Completed, under your role, with every change in the board’s History under the assistant’s name.

For parallel runs the useful tool is fenbs_next_approved_task: it hands an assistant the most urgent task a person has approved for AI and holds it for that assistant while it works, so two runs never pick up the same one. fenbs_release_task gives it back if the run cannot finish. fenbs has no sprints, due dates or assignee field to manage; the lanes are To Do, Next Up, In Progress and Completed. Setup is on Codex CLI on fenbs.

Related

Every command and config key in one place: Codex CLI commands. The same questions for another vendor: Claude Code subagents vs agent teams and multi-agent workflows. Another terminal agent: Cursor CLI. The tools behind the board: the MCP docs.

Questions people ask.

Can Codex CLI spawn subagents?

Yes. Current Codex releases enable subagent workflows by default in the CLI, the IDE extension and the desktop app. At most intelligence levels Codex spawns them only when you ask, or when AGENTS.md or a skill tells it to. Each subagent works in its own thread and returns a summary.

Where do Codex custom agents go?

In TOML files, one agent per file: ~/.codex/agents/ for your own and .codex/agents/ in a project. Each needs name, description and developer_instructions, and can set model, reasoning effort and sandbox mode.

How many subagents can Codex run at once?

You set the cap with max_concurrent_threads_per_session in the [agents] table of config.toml. If you leave it unset, Codex chooses the default, and OpenAI’s documentation does not state the number.

Can I still run Codex as an MCP server for the Agents SDK?

No. OpenAI removed the codex mcp-server command and the standalone binary. Integrations now use the Codex app server or the Codex SDK, and Codex remains an MCP client for other servers.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.