Context Engineering for AI Agents: A Practical Guide
An agent only knows what is in its context window at the moment it acts. Context engineering is deciding what goes in, what stays out, and where the rest lives so the next session can find it.
7 min read
Context engineering for AI agents is the work of choosing what sits in the model’s context window at each step: the standing instructions, the tool definitions, the files and results it has fetched, the conversation so far, and whatever it remembers from before. Anthropic’s engineers describe it as the strategies for curating and maintaining the optimal set of tokens (opens in a new tab) during inference. In practice it comes down to five habits: keep instructions short and stable, fetch information when it is needed rather than up front, give the agent few tools with clear jobs, summarise history before it crowds out the work, and keep anything that must last somewhere outside the window.
What is actually in the window
Before you can manage context you need an inventory. For a coding agent such as Claude Code, a typical turn carries all of these at once:
- A system prompt you never see, plus your own instructions file (
CLAUDE.md,AGENTS.mdor the equivalent for your tool). - The names and descriptions of every tool it can call, including those from connected MCP servers.
- Anything it has read: files, search results, web pages, command output.
- The conversation itself: your messages, its replies, and every tool call and result in between.
- Memory: notes carried over from earlier sessions.
Each of these competes for the same space. Every file read and every long test log stays in the window until something removes it, and the model attends to all of it when it decides what to do next. The rest of this guide takes each item in turn.
Instructions: short, stable, specific
The instructions file is the one part of context you write once and pay for on every turn. That argues for keeping it small. Anthropic’s advice is to aim for the “right altitude”: specific enough to steer behaviour, general enough to give the model heuristics rather than a brittle script. Claude Code’s own documentation on how Claude remembers your project (opens in a new tab) suggests keeping each CLAUDE.md under 200 lines, because longer files use more context and reduce adherence.
What belongs in it: commands the agent cannot guess, rules that differ from the defaults, things it must never touch, and where things are. What does not: anything it can work out by reading the code, and anything that changes weekly. How to lay the file out across several repositories is its own topic, covered in managing multiple projects with CLAUDE.md and Claude Code project structure.
Retrieval: just in time, not all at once
The tempting approach is to load everything relevant before the agent starts: the whole spec, every related file, the last month of tickets. The alternative is to give it pointers — file paths, a search command, a task reference — and let it fetch what it needs when it needs it. Anthropic calls this just-in-time context, and describes Claude Code as a hybrid: CLAUDE.md goes in up front, and tools like glob and grep let it find files as the work demands.
- Name the place, not the contents: “the retry logic is in
src/payments/retry.ts” costs a line; pasting the file costs hundreds. - Prefer a search the agent can run over a list you maintain by hand. Lists go stale;
grepdoes not. - Put rarely needed knowledge where it loads on demand. In Claude Code that means skills or path-scoped rules rather than the root instructions file.
Tools: fewer, with clear jobs
Every tool definition sits in the window, and every overlapping pair is a decision the model can get wrong. Anthropic’s test is a good one: if a human engineer cannot say for certain which tool fits a situation, the agent will not do better. Connect the MCP servers the work needs, not every one you have, and prefer tools that return a short, relevant answer over tools that dump a whole record.
The same goes for tool output. A test run that prints ten thousand passing lines puts ten thousand lines in context. Ask for the failures and the summary line instead.
History: compact before it crowds the work
Long tasks outgrow the window. Compaction is the standard answer: summarise the conversation so far and carry on from the summary. Claude Code does this automatically as the window fills, and /compact lets you do it yourself with a focus, for example /compact keep the plan and the list of files changed. Between unrelated jobs, /clear starts over.
The useful detail is what survives compaction (opens in a new tab). The project-root CLAUDE.md, auto memory and a plan written in plan mode are re-injected from disk; the conversation, and anything the agent only learned in it, is summarised. So a decision you need later should be written to a file or a task before the window is compacted, not left in chat. Why long windows degrade in the first place is covered in context rot.
Memory and delegation
Two techniques keep a long job coherent without keeping everything in one window. The first is structured note-taking: the agent writes progress, open questions and decisions to a file outside the window and reads them back later. Claude Code’s auto memory is a built-in version, loading the first 200 lines or 25KB of its MEMORY.md at the start of each session.
The second is delegation. A subagent (opens in a new tab) does its reading in a separate window and hands back a summary, so the main conversation receives the conclusion rather than the dozens of files behind it. When that is worth it, and how it relates to the Task tool, is in Task tool vs subagents.
A task board as context that outlives the session
Notes files and auto memory belong to one machine and, usually, one assistant. The work itself often does not: a job started in Claude Code on Monday may be finished by Cursor on Wednesday, or by a colleague. For that you want state kept with the work rather than with the tool, and a task board over MCP is a natural place for it.
On a fenbs board, each task carries a note (what the problem is, why, where), a plan (how it will be done, rewritten as the agent learns), test notes, and a thread of comments. The board also keeps AI context: short notes that every connected assistant reads on arrival with fenbs_get_context, either the notes for everything or those plus one project’s own. An assistant that works something out can leave a note with fenbs_add_context_note, signed with its name, for the next one.
fenbs_whoami # which board, which role fenbs_get_context project: "api" # the notes every assistant should read first fenbs_list_items board: "acme", lane: "doing" # what is already under way fenbs_get_item ref: "BUG-031" # the note, the plan and the comments so far
Four calls and the new session knows what it is working on, what has been tried and what the people on the board want every assistant to know, without inheriting the previous session’s transcript. That is the point of persistent context: carry the conclusions forward and leave the noise behind.
A checklist to start from
- Read your instructions file line by line and cut anything the agent could learn from the code.
- Replace pasted content with pointers: paths, search commands, task references.
- Disconnect the tools and MCP servers this project does not use.
- Tell the agent what to keep when compacting, and write decisions to a file or a task, not only to chat.
- Send broad investigations to a subagent and keep the main window for the change itself.
- Keep the state of the work — plan, progress, blockers — somewhere the next session and the next assistant can read.
Related
For the day-to-day rhythm of reading and updating a board from Claude Code, see a task-tracking workflow for Claude Code. To connect an assistant, start with the MCP docs or the Claude Code integration.