What Is an Agent Harness? Why the Loop Around the Model Matters

The harness is everything in an agent that is not the model: the loop, the tools, what goes into the context, the permissions, the memory and the sandbox. Why the same model does better in one harness than another, what Claude Code, Codex, VS Code and the SDKs mean by the word, and what it changes when you choose tools.

7 min read

An AI agent harness is everything in an agent except the model. The model decides what to do next; the harness runs the loop that asks it, supplies the tools it can call, decides what goes into its context and what gets compacted out, enforces what it may do without asking, keeps memory between sessions, and fences commands inside a sandbox. Claude Code, Codex and the agents inside VS Code are harnesses around someone’s model, and so is the agent you build with an SDK. The harness matters because it changes results: the same model, in two harnesses, can finish a task in one and wander in the other. When you choose a coding agent you are mostly choosing a harness.

Where the word comes from

Anthropic uses it for its own product. Its page on how Claude Code works (opens in a new tab) describes an agentic loop powered by models that reason and tools that act, and says Claude Code is the layer around the model that provides the tools and manages the context the model sees, which is what the term agentic harness refers to. Anthropic also calls the Claude Agent SDK a general-purpose agent harness.

LangChain’s write-up on the anatomy of an agent harness (opens in a new tab) put it as a formula: agent = model + harness, and if you are not the model, you are the harness. Its list of what a harness holds is a useful checklist: system prompts; tools, skills and MCP servers with their descriptions; a filesystem, a sandbox and a browser; orchestration such as subagents, handoffs and model routing; and hooks or middleware for the steps that must happen every time.

What a harness does

  • The loop: call the model, run the tools it asks for, append the results, and decide when to stop, including turn limits and what happens after an error.
  • Tools: which exist at all, how they are described, and how their output is trimmed before the model sees it.
  • Context: the system prompt, instruction files such as AGENTS.md or CLAUDE.md, which tool definitions load up front and which wait, and how an overflowing conversation is compacted.
  • Permissions: what runs without asking, what asks, and what is refused, from per-command allowlists to classifiers that review risky calls.
  • Memory: what survives the end of a session, from memory files to progress notes the next session reads first.
  • Sandboxing: where commands run, what the filesystem and network look like from inside, and how changes are isolated, for example in a worktree.
  • Verification: running the tests, reading type errors after an edit, and checking the result before calling it done.

Why the same model behaves differently

Each item above changes what the model sees and what happens when it acts, so each one changes the outcome. Some examples of the difference:

  • The system prompt. The Claude Agent SDK starts with a minimal prompt unless you ask for Claude Code’s preset, so the same model is a different agent with and without it.
  • Tool definitions. Claude Code defers MCP tool definitions by default and loads them through tool search, so only names take up space until a tool is used. A harness that loads every definition up front leaves less room for the work.
  • Compaction. What a harness keeps when the context fills, and what it drops, decides whether the agent still remembers the instruction you gave at the start.
  • Verification. A harness that runs the tests after each edit, and feeds failures back, turns a guess into a checked change.
  • Tool reach. An assistant in VS Code’s Copilot harness can reach fewer MCP servers than one in the Local harness of the same editor, because the Copilot harness supports fewer kinds of server.

LangChain gives a measured example of the size of the effect: by changing only the harness around the same model, it says it moved its coding agent from outside the top 30 to the top 5 on the Terminal Bench 2.0 benchmark. Treat the figures as one team’s result on one benchmark, not a rule.

Harnesses you already use

  • Claude Code: one agentic loop behind the terminal, the IDE extensions, the desktop app and the web, with checkpoints, permission modes, compaction, CLAUDE.md, skills, hooks and subagents. The Claude Agent SDK is the same harness as a library; see the Claude Agent SDK.
  • Codex: OpenAI’s write-up on Codex as a platform (opens in a new tab) says the harness manages conversation state, streams execution, uses tools, enforces sandbox and approval policies and carries work across turns. The app, CLI and IDE extension share it, it is open source, and you can drive it with codex exec, the Codex SDK or the app-server. OpenAI’s Agents API also runs agents on the Codex harness.
  • VS Code: the editor now runs several harnesses side by side. Its page on agent harnesses (opens in a new tab) lists Local, Copilot, Claude and Codex, plus a Cloud target. The Copilot harness can currently reach only local MCP servers that do not require authentication, and Cloud sessions use the tools and MCP servers the cloud service provides, not your editor’s. Which configuration each reads is in the .vscode/mcp.json guide.
  • Microsoft Agent Framework: it ships a harness as a product. Its Harness agent (opens in a new tab) defines a harness as the runtime scaffolding that turns a model into an agent, and adds todo tracking, plan and execute modes, compaction, file memory and standing tool approvals by default. More in Microsoft Agent Framework.
  • Frameworks in general: LangGraph, CrewAI, the OpenAI Agents SDK and Google ADK give you the parts to build a harness of your own. AI agent frameworks compared sorts them.

Harness engineering

The practice got its name in 2026. OpenAI’s engineering blog published “Harness engineering: leveraging Codex in an agent-first world”, on building software where agents write the code and engineers design the environment, the instructions and the feedback loops that keep them on track. Birgitta Böckeler’s article on harness engineering for coding agents (opens in a new tab) on martinfowler.com sorts the controls into guides, which steer an agent before it acts, such as instruction files and skills, and sensors, which check afterwards and let it correct itself, such as tests, linters and reviews. Some are computational and deterministic; others use a model to judge.

Anthropic’s engineering post on harnesses for long-running agents is a worked example. Because each new session starts with no memory, its harness has a first session write a feature list and a progress file, and every later session read them and the git log, work on one feature, commit, and update the notes. None of that is the model. All of it is harness.

What it means for choosing tools

  1. Compare harnesses, not only models. Most coding agents let you pick among several models, so the model is often the easier thing to change. Ask what each harness does with permissions, context, verification and sandboxing.
  2. Test on your own work. Run the same real task in two agents with the same model. The difference you see is the harness.
  3. Check what the harness can reach. MCP support differs by harness even inside one editor, as VS Code’s Copilot harness shows.
  4. Invest in the parts that move between harnesses: an AGENTS.md most agents read, skills in the open format, MCP servers, and tests. They are your harness engineering, and they survive a change of tool.
  5. For a team, agree one set of guides and sensors and let people pick their agent. The best AI coding agents compares the agents themselves.

The part of the harness that no vendor owns

Two harness jobs work better outside any one agent: remembering what is to do and what was done, and enforcing what an assistant may change. A shared board does both. On fenbs, every assistant, whatever harness it runs in, reaches the same tasks over MCP and reads the same AI context before it starts, and the next session picks up the task where the last one commented. The limit is enforced by the board rather than the harness: each connection carries the scopes ticked when it was approved, checked against the person’s role, so a harness set to approve everything still cannot move a task its connection may not move. Every change is recorded in the board’s history under the assistant’s name.

Related

Choosing an agent: the best AI coding agents. Building your own harness: the Claude Agent SDK. What goes into the context, and why it matters: context engineering for AI agents. The notes every assistant reads first: AI context.

Questions people ask.

What is an agent harness?

It is everything in an AI agent except the model: the loop that calls the model and runs its tools, the tool set, what goes into the context and how it is compacted, the permission rules, memory between sessions and the sandbox commands run in. Claude Code, Codex and VS Code’s agents are harnesses.

What is the difference between an agent harness and an agent framework?

A framework is a library of parts for building agents. A harness is a complete runtime around a model, ready to work. Claude Code and Codex are harnesses; LangGraph and the OpenAI Agents SDK are frameworks you could build one with. Some products are both, such as the Claude Agent SDK and Microsoft Agent Framework’s Harness agent.

What is harness engineering?

The practice of designing everything around a coding agent so it does reliable work: instruction files, skills, tools, tests, linters and review steps that steer it beforehand and check it afterwards. The term was used in 2026 by OpenAI’s engineering blog and taken up by LangChain, Martin Fowler’s site and others.

Why does the same model perform differently in different tools?

Because each tool is a different harness. The system prompt, the tool descriptions, how context is compacted, which tools load, whether tests are run after edits and what is allowed without asking all change what the model sees and what happens when it acts.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.