AI Context Windows Explained for People Who Manage Work
An AI assistant can only work with what is in its context window at that moment. What a context window is, what fills it, why a bigger one is not the whole answer, and what that means when you hand work to an assistant.
7 min read
An AI context window is everything a model can take into account when it produces its next answer: your instructions, the files and documents it has been given, the definitions of the tools it can use, the conversation so far, and the answer it is writing. It is measured in tokens, it has a fixed limit, and it is empty at the start of every new session. For anyone handing work to an AI assistant, that last point matters most. The assistant does not remember your project the way a colleague does; it knows what is in the window right now, and nothing else, unless something puts it back.
Working memory, not a filing cabinet
Anthropic’s glossary (opens in a new tab) describes the context window as a “working memory” for the model, and separates it from the large body of text the model was trained on. Training is where the model learned language, facts and skills in general. The context window is where your particular job lives: this spreadsheet, this bug, this client’s brief.
A useful picture is a desk. The model is well read, but it can only work on the papers currently on the desk. The desk has a fixed size. When the session ends the desk is cleared, and Claude Code’s own memory documentation opens with exactly that: each session begins with a fresh context window (opens in a new tab). Anything that should carry over has to be written somewhere and put back on the desk next time. How assistants do that is the subject of agent memory.
Tokens, briefly
Windows are measured in tokens rather than words. A token is a chunk of text: sometimes a whole word, often part of one, sometimes a single character or punctuation mark. For Claude, the glossary says a token is roughly 3.5 English characters on average, varying by language. Other models split text differently, so the same document can be a different number of tokens on different assistants.
You rarely need to count them, but a sense of scale helps. In a post on retrieval for large knowledge bases (opens in a new tab), Anthropic equates 200,000 tokens with about 500 pages of material. Claude’s current models range from windows of that size up to a million tokens, depending on the model. Those are large numbers, and it is easy to conclude the limit no longer matters. The next two sections explain why it still does.
What fills the window
The Claude API documentation on context windows (opens in a new tab) is precise about what counts: the system prompt, every message including tool results, images and documents, the tool definitions, and the output the model generates, including its thinking. In an everyday session with an assistant, that turns into five kinds of material:
- Instructions. The assistant’s built-in system prompt, which you never see, plus any standing instructions you have given it, such as a project file like
CLAUDE.mdor custom instructions in a chat app. - Tool definitions. Every tool the assistant can call has a name and a description sitting in the window, including those from connected MCP servers. Connect a lot of tools and they take up room before any work begins.
- Things it has read. Files, spreadsheets, web pages, search results, command output, and anything you paste in.
- The conversation. Your messages, its replies, and every tool call and result in between. Nothing drops out on its own until the limit is near.
- Its own output. The answer it is writing and any reasoning it does first.
A lot of this arrives before you type anything. Claude Code’s interactive guide to the context window (opens in a new tab) walks through a session: the system prompt, auto memory, environment details, the names of MCP tools, skill descriptions and CLAUDE.md files all load first, and your opening request is small beside them. The same guide shows the running total growing with every file read, which is why a long session feels heavier than a short one.
Why size is not everything
A bigger window means more fits before the assistant runs out of room. It does not mean the assistant uses everything in it equally well. Anthropic’s documentation says plainly that more context is not automatically better, and that accuracy and recall degrade as the token count grows; the gradual version of that effect is covered in context rot.
There are practical costs as well. The whole window goes back to the model on each turn, so a stuffed window is slower to answer and, where usage is metered, dearer. And when the limit does arrive, something has to give. Assistants handle that in different ways: chat apps may drop the oldest turns, and coding agents such as Claude Code summarise the conversation so far, a step called compaction. Either way, detail that lived only in the conversation can be lost.
So the useful question is not how big the window is, but how much of what is in it bears on the job at hand.
What it means when you hand work to an assistant
If you manage work rather than write code, the context window explains most of the surprises people have with AI assistants. It explains why an assistant that was brilliant yesterday asks today what the project is, why it forgets a rule from the morning by late afternoon, and why it can confidently act on a decision you reversed an hour ago but only mentioned once. Five habits follow from it.
- Brief it every time, and briefly. Assume a new session knows nothing about the work. Give it what it needs for this job, not the history of the whole project.
- Point, do not paste. “The pricing rules are in the Q3 sheet, tab Rates” costs a line; pasting the sheet costs thousands of tokens, most of them irrelevant to the question.
- One job per session. Mixing three unrelated tasks in one long chat fills the window with material that has nothing to do with the task at hand.
- Write decisions down outside the chat. If a decision only exists in the conversation, it is one summary away from being lost. Put it where the next session will read it.
- Connect only the tools the job needs. Each connection adds definitions to every turn, and more choices mean more ways to pick the wrong one.
The common thread is that the brief, the decisions and the progress of the work should live somewhere other than the assistant’s window. A task is a natural place for them. On a fenbs board each task holds a note (what the problem is, why, where), a plan (how it will be done, rewritten as the work goes on), and a comment thread, and an assistant connected over MCP fetches exactly that task when it starts. Standing guidance that applies to everything, such as naming conventions or what never to touch, lives in the board’s AI context, which each assistant reads on arrival.
fenbs_get_context project: "website" # standing notes for this project fenbs_get_item ref: "FET-042" # the note, the plan and the comments so far
Two short calls, and the new session has the brief without inheriting the previous session’s thousands of lines of tool output.
Questions to ask about any assistant
- What loads automatically at the start of a session, and can I see it? In Claude Code,
/contextshows what has loaded into the current session. - What happens when the window fills: are old turns dropped, or summarised, and what survives?
- Where does it keep things between sessions, and can the people I work with read them?
- How many tools are connected, and does each one earn its place?
Related
For the working discipline behind all of this, read context engineering for AI agents. For writing the brief itself, see how to write a task for an AI agent. To connect an assistant to a board, start with the MCP docs.