Context Engineering for Multi-Agent Systems: What Each Agent Should See
Splitting work across agents only helps if each one sees the right things and nothing else. How to write the brief, isolate each agent’s context, pass back summaries instead of transcripts, keep shared state outside the agents, and stop instructions contradicting each other.
7 min read
Context engineering for multi-agent systems is deciding what goes into each agent’s context window, not just the main one. Each agent should see the standing rules everyone shares, the instructions for its own role, a brief written for its piece of work, and the tools that piece needs. It should not see the other agents’ transcripts. What comes back from it should be a short result in an agreed shape, with pointers to anything longer. And anything more than one agent needs to know, such as the plan, the progress and the decisions, belongs outside all of them, where every agent reads the same copy.
How to split the work and keep agents from taking the same task is covered in multi-agent workflows. The single-agent version of this subject is context engineering for AI agents. This post is about the context each agent in the system gets.
Why several agents is a context decision
The main reason to use more than one agent is that each gets a clean window. Describing its multi-agent research system (opens in a new tab), Anthropic says subagents facilitate compression by working in parallel with their own context windows, and give separation of concerns: distinct tools, prompts and exploration paths. The same write-up gives the price: in its data, multi-agent systems used about fifteen times more tokens than chats. You are buying focus with tokens, so it is only worth it if each agent’s window really is more focused than one agent’s would have been.
That gives a test for every design choice below. If an agent ends up carrying the whole conversation, every other agent’s output and every tool the system has, you have paid for several windows and got several copies of one.
The four layers each agent sees
- Standing rules: what is true for every agent on the project, such as build commands, never-touch areas and naming. The same file for all of them.
- Role instructions: what makes a reviewer a reviewer. Short, and only about the role.
- The brief: the piece of work, written by whoever delegates it. This is the part that changes every time.
- Tools: only those the role needs. Every tool description is text the model reads, and every unnecessary one is a choice it can get wrong.
Notice what is not on the list: the lead agent’s conversation, the other workers’ results, and the full history of the job. Leave those out by default and add a piece back only when a worker cannot do its job without it.
Write the brief as if the agent knows nothing
In most tools a delegated agent starts without the conversation that led to it. Claude Code’s subagents documentation (opens in a new tab) says a subagent does not see your conversation history, and VS Code says the same of its local subagents. So the brief carries everything. Anthropic found that vague briefs lead agents to duplicate work, leave gaps or miss information, and that each subagent needs an objective, an output format, guidance on the tools and sources to use, and clear task boundaries.
Objective: find every call site that writes to the orders table without going through OrderRepository, in services/ and jobs/. Why: we are adding an audit column; direct writes would skip it. Boundaries: read only. Do not edit files. Ignore tests/ and scripts/. Tools: code search and file read. No web access. Return, in this shape and nothing else: - file:line and one line on what the write does, per call site - anything you were unsure about, as a question - under 300 words in total
The “Why” line is the one people drop. Without it the worker cannot judge an edge case the brief did not foresee, and will either guess or report everything.
Isolate context per agent, on purpose
Frameworks differ in what they pass along by default, and the default is not always what you want. In the OpenAI Agents SDK, a handoff (opens in a new tab) works as though the new agent takes over the conversation, and it sees the entire previous history unless you add an input filter; the SDK ships one that removes tool calls and results from that history. Its alternative is a manager that calls specialists as tools and keeps the conversation itself. Claude Code subagents and VS Code local subagents start clean and receive only the task they are given.
- Use a clean start for workers doing a bounded job: a search, a review, a test run. The brief is their whole world.
- Use a handoff with history when the next agent must answer the same person about the same thread, such as a triage agent passing a customer to a specialist. Filter out the tool noise.
- Give each agent its own tool list. A reviewer with no edit tools cannot edit, whatever its instructions say.
- Keep isolated context and isolated files apart in your head. A worker with its own window can still write over another’s files; separate checkouts are a different control.
Pass summaries back, not transcripts
What comes back from a worker lands in the lead’s window, so its size and shape are a context decision. Anthropic’s article on effective context engineering (opens in a new tab) describes subagents that do extensive work but return a condensed, distilled summary, often 1,000 to 2,000 tokens, so the detailed search context stays isolated in the subagent and the lead can focus on synthesis.
- Fix the return format in the brief: a list, a table, a verdict and three reasons. A lead combining five results needs them in the same shape.
- Return pointers, not contents:
file:line, a commit hash, a task reference, a path to a report. The lead can open one if it needs to. - Put long outputs somewhere they persist. Anthropic recommends letting agents write outputs that exist independently of the lead rather than routing everything through it, which avoids losing detail in the retelling.
- Ask for uncertainty explicitly. A summary that sounds confident about everything hides the one thing the lead should have checked.
Keep shared state outside the agents
Some information every agent needs: the overall plan, what is done, what is blocked, what was decided. If it lives in the lead’s window, it is at the mercy of that window. Anthropic’s lead researcher saves its plan to memory at the start because a context over 200,000 tokens is truncated, and the plan must survive that. The general form is a store outside every agent that each reads when it starts and writes when it stops.
A task board reached over MCP is one such store, and it has the advantage of being readable by people and by agents in different tools. On a fenbs board, each task keeps the problem in its note and the approach in its plan, which an agent rewrites as it learns, plus a testing status and a thread of comments. Workers read their task with fenbs_get_item instead of receiving the lead’s memory of it, and write their result back as a comment. Related tasks link both ways with relatesTo, so a follow-up stage can find where it came from. Everything any agent changes is recorded in the board’s history with its name, which is how you reconstruct afterwards what each one did.
Stop instructions from contradicting each other
More agents means more places for instructions to live: the shared rules file, each role file, each brief, each tool’s description. When two of them disagree, the model picks one, and different agents may pick differently. That is how a system produces two workers following two versions of the same rule.
- Write each rule once. Standing rules live in one shared file every agent reads; role files add to it and never restate it.
- Briefs describe the job, not the rules. If a brief needs to override a standing rule, change the rule or say explicitly that this job is an exception.
- Record decisions where every agent can check them. On fenbs that is the Decisions page:
fenbs_list_decisionsshows what people have decided, and an assistant can ask an open question withfenbs_add_decisionbut never decides one itself. - Leave notes for the next agent in one place.
fenbs_get_contextgives every assistant the same AI context notes on arrival, andfenbs_add_context_noteadds one; update an existing note rather than adding a second that contradicts it.
A checklist before you add an agent
- Can one agent with a subagent for the heavy reading do this? If so, start there.
- For each agent, list what it sees: rules, role, brief, tools. Remove anything it does not need.
- Write the return format into every brief, with a length limit.
- Decide where shared state lives before the first agent runs, and tell every agent to read it first.
- Read the standing rules and every role file side by side, looking for two lines that disagree.
Related
Coordinating who takes which task: multi-agent workflows. The two ways Claude Code splits work: subagents vs agent teams. Passing work between agents and people: handing work between AI agents and people. Connect an assistant to a board with the MCP docs.