Why MCP Servers Eat Your Context Window, and What to Do
Every MCP server you connect adds tool definitions to the model’s context, and every result it returns stays there. Where the tokens go, how to see them in each client, and the fixes that work: tool search, trimming, toolsets, leaner servers and code execution.
8 min read
MCP servers use tokens in two places. The first is the tool list: each tool’s name, description and input schema is sent to the model as part of its context, on every request, whether or not the tool is called, unless your client defers it. The second is results: whatever a tool returns is added to the conversation and re-sent with every later message. Anthropic’s tool search documentation (opens in a new tab) puts a typical five-server setup at around 55,000 tokens of definitions before any work starts, and says tool choice gets worse beyond 30 to 50 tools. The fixes, in order of effort: use a client that loads tool definitions on demand, switch off servers and tools you do not need, narrow big servers with their toolset options, ask for less in each call, and, for heavy workloads, let the agent call MCP through code instead of directly.
This page covers MCP across clients. The Claude Code commands and settings, including /usage, /context and /mcp, are in how to reduce Claude Code token usage; the wider discipline of deciding what sits in the window is context engineering for AI agents.
Where the tokens go
- Definitions. A tool is a name, a description and a JSON Schema for its arguments, sometimes an output schema too. A server with 40 tools and careful descriptions can run to tens of thousands of characters. Prompt caching makes re-sending them cheaper, not free, and they still take room the model could use for your code or documents.
- Results. A tool that returns a whole record when you needed one field, or 500 rows when you needed ten, puts all of it in the conversation. It stays there until the conversation is cleared or compacted, and it is paid for again on each later turn.
- Round trips. When the agent reads data with one tool and passes it to another, the data travels through the model twice: once as a result, once as the next call’s arguments.
- Wrong turns. More tools mean more near-misses. A model that picks the wrong one of three similar tools spends a call and a result finding out.
Measuring it, client by client
- Claude Code:
/contextshows how much of the window MCP tools take, and/usagebreaks recent usage down by MCP server on subscription plans. The details are in the Claude Code post linked above. - VS Code: a chat request can carry at most 128 tools, according to VS Code’s page on using tools in chat (opens in a new tab). If you hit it, deselect tools or whole servers in the tools picker, or turn on the
github.copilot.chat.virtualTools.thresholdsetting, which groups tools automatically. - Cursor: servers can be toggled off without deleting them, and it loads MCP tools on demand (below). A per-server token count is not documented.
- Codex:
enabled_toolsanddisabled_toolsinconfig.tomllimit what each server offers. A per-server token count is not documented. - Any client: call
tools/listyourself with the MCP Inspector and look at the size of the response. That JSON, give or take formatting, is what a client without deferred loading adds to every request.
Tool search and deferred loading
The biggest fix is on the client side. Instead of sending every definition, the client sends only tool names, or a search tool, and the model looks up the few definitions it needs when it needs them. Anthropic’s API does this with defer_loading: true on each tool and a tool search tool, and for MCP servers reached through its MCP connector the setting goes once on the server’s toolset. The same documentation recommends tool search once you have ten or more tools or more than 10,000 tokens of definitions, and plain tool calling below that.
- Claude Code: tool search is on by default. Only tool names and server instructions load at the start, full definitions are fetched on demand, and a server you need every turn can be exempted with
alwaysLoad. - Cursor: MCP tools are loaded on demand. Cursor’s post on dynamic context discovery (opens in a new tab) says it syncs tool descriptions to a folder, the agent receives the tool names and looks up the rest when needed, and in an A/B test of runs that called an MCP tool this cut total agent tokens by 46.9%, with high variance depending on how many servers were installed.
- VS Code: the virtual tools setting above groups tools once the count passes a threshold, and the model activates a group before calling its tools.
Deferred loading has a cost: a search step, and a chance that the model does not think to look for a tool it cannot see. Server instructions and clear, keyword-rich descriptions are what make a deferred tool findable, which is why they matter more, not less, once tool search is on.
Trim what you connect
- Connect servers per project, not globally. A database server belongs in the one repository that uses it, not in every session you start.
- Switch off servers you are not using this week. Every client above can disable a server without deleting its configuration.
- Remove duplicates. Two servers for the same service, say an official one and a community wrapper, double the definitions and give the model two near-identical choices.
- Prefer a command-line tool when one does the job. A CLI the agent already knows adds no tool listing at all; MCP vs CLI sets out when each is the better fit.
Toolsets and headers
Large vendor servers let you choose which groups of tools to expose. The GitHub MCP server’s README (opens in a new tab) says enabling only the toolsets you need helps the model choose tools and reduces the context size. On the hosted server you pick them with the X-MCP-Toolsets header, or name single tools with X-MCP-Tools; left out, you get the defaults: context, repositories, issues, pull requests and users. Azure DevOps’s remote server takes the same X-MCP-Toolsets header with groups such as repos, wit and pipelines. A read-only header removes write tools as well, which shrinks the list further.
{
"mcpServers": {
"github": {
"type": "http",
"url": "https://api.githubcopilot.com/mcp/",
"headers": {
"Authorization": "Bearer ${GITHUB_PAT}",
"X-MCP-Toolsets": "issues,pull_requests",
"X-MCP-Readonly": "true"
}
}
}
}The full set of GitHub options is in Claude and the GitHub MCP server, and the Azure DevOps toolsets and domains in Azure DevOps MCP server.
Server-side design that cuts tokens
If you build servers, most of the saving is yours to make. Anthropic’s guide to writing tools for agents (opens in a new tab) is the best single reference; the points that matter most for tokens:
- Fewer, broader tools. One
schedule_eventtool that does the lookups itself beats separate tools to list users, list events and create an event, because the model reads one definition and makes one call. - Filters, pagination and sensible defaults. Let the caller ask for one project, one status or the first 20 items, and default to a small page.
- A choice of detail. The guide’s example returns the same result in 206 tokens in detailed form and 72 in concise form, chosen by a
response_formatargument. - Meaningful fields only. Names and titles rather than internal IDs and MIME types, and no fields the model will never use.
- Short descriptions with the critical part first. Clients may truncate long ones, and every word is re-sent on every request where definitions are loaded.
- Server instructions that say what the server is for, so a client using tool search knows when to look.
Tool design in general, with a worked server, is covered in how to build an MCP server.
Code execution with MCP
The largest reductions come from changing who handles the data. In code execution with MCP (opens in a new tab), Anthropic describes presenting MCP servers to the agent as code: a file tree such as servers/google-drive/getDocument.ts, one file per tool. The agent lists the folder, reads only the definitions it needs, and writes a short program that calls the tools and passes data between them inside a sandbox. Large intermediate results never enter the model’s context; only what the program prints does. In Anthropic’s example, reading a meeting transcript from Google Drive and adding it to a Salesforce record fell from about 150,000 tokens to about 2,000.
The same post is clear about the price: running code the agent wrote needs a secure execution environment with sandboxing, resource limits and monitoring. It pays off when an agent moves a lot of data between tools or uses many servers; for a handful of small calls, direct tool calls are simpler and fine.
What one real server costs
fenbs, a task board where people and AI assistants are members with roles, offers 30 tools at https://fenbs.ai/api/mcp. Its full tool list, measured from the server’s own definitions, is about 25,500 characters of JSON. Four tools, for adding and changing tasks and decisions, make up nearly half of it, because their descriptions carry the rules an assistant should follow, such as keeping the problem and the plan in separate fields. In a client with deferred loading, an idle fenbs connection costs little more than its tool names; without it, it is a fixed charge on every request, which is the case for switching it off in projects that do not use the board. On results, fenbs_list_items takes filters for lane, kind, project, category, free text and flag, so asking for the In Progress lane of one project returns those tasks and nothing else.
Related
Connecting an assistant and the full tool list: MCP docs. Why a crowded window also makes answers worse: AI context rot. How big the window is in the first place: AI context window. What fenbs keeps between sessions so assistants re-read less: AI context.