Agent Memory vs Context: What Each Holds
Context is what the model can see in this one call, and it ends with the call. Memory is what is kept somewhere else and brought back when it is needed. How the two meet, three everyday examples, and a way to decide what belongs in each.
7 min read
Context is everything the model can see while it produces one answer: the instructions, the conversation so far, the tool results, the files it was handed. It is bounded by the context window and it is gone when the call ends. Memory is whatever is kept outside the model — a file, a database, a vector store, a task board — and selectively put back into context on a later call. An agent only ever reasons over context. Memory matters only at the moment some of it becomes context again.
The difference at a glance
- Where it lives: context is inside the request; memory is in storage you or a tool control.
- How long it lasts: context lasts one call (or one conversation, if the history is sent again each turn); memory lasts until someone deletes it.
- Size: context is capped by the model’s window; memory is capped by nothing but good sense.
- Who chooses what is in it: context is assembled for each call; memory is written when something is worth keeping.
- What goes wrong: context overflows or gets noisy; memory goes stale or never gets read.
Context: what the model sees in this call
Anthropic’s documentation on context windows (opens in a new tab) describes the window as the model’s “working memory”: all the text it can reference when generating a response, including the response itself. Everything in the request counts towards it — the system prompt, every message, every tool result and document, and the tool definitions — and so does the output.
Two properties follow. First, a conversation only feels continuous because each turn sends the history again; the model does not hold it between calls. Second, filling the window is not free even when it fits. The same page notes that accuracy and recall degrade as the token count grows, which is why curating what goes in matters as much as how much room there is. Our explainer on the AI context window and the post on context rot go further into both.
Memory: what persists and comes back
Memory is a store plus a way back in. Anthropic’s memory tool (opens in a new tab) is a clean example of the shape: Claude asks to create, read or edit files under a memory directory, your application carries out those operations against storage you choose, and a later conversation reads the files back. The documentation calls this just-in-time retrieval — the agent records what it learns and loads it on demand instead of carrying everything in the window.
That last part is the point. Memory that is never read back has no effect on the next answer, however carefully it was written. So every memory design is really two decisions: what to write down, and how the right piece gets back into context at the right time.
How memory becomes context
There are four common ways back in, and they differ mostly in how much context they spend and who decides what is loaded.
- Loaded whole at the start. A rules file is read into every session. Simple and reliable, but every line costs context on every call.
- An index loaded, detail on demand. A short summary is always present; the agent opens the longer notes only when the summary says they are relevant.
- Retrieved by similarity. A vector store returns the passages that look most like the current question. Cheap on context, but only as good as the match.
- Queried through a tool. The agent calls a tool — a search, a board, a database — and only the answer enters context.
Three examples
CLAUDE.md: memory that is always context
Claude Code’s memory documentation (opens in a new tab) says each session begins with a fresh context window, and two mechanisms carry knowledge across: CLAUDE.md files you write, and auto memory Claude writes itself. Both are loaded at the start of every conversation — auto memory only up to the first 200 lines or 25KB of its MEMORY.md index, with topic files read when needed. The same page is candid that Claude treats these as context, not enforced configuration. That is the first pattern and the second pattern side by side, and the trade-off is the one above: what is always loaded is always paid for. The post on Claude Code memory covers the details.
A task board: memory reached through a tool
A board holds what is happening: which tasks exist, which lane each is in, what was tried and what was decided. None of that belongs in the window by default. The agent asks for it when it needs it — the open bugs in one project, the comments on one task — and only that answer is loaded. The board also has something a private memory file lacks: people read and change the same records, and each change carries a name.
A vector store: memory reached by resemblance
A vector store is the usual answer when there is too much to load and no obvious index: past tickets, meeting notes, long documents. It is good at “find something like this” and poor at “what is the current state”, because the most similar passage is not necessarily the newest. Agent memory vs RAG looks at that distinction on its own.
The same split in frameworks
Frameworks often call both of these memory, which is where the confusion starts. LangGraph’s memory concepts page (opens in a new tab) splits it into short-term memory, which is thread-scoped and tracks the ongoing conversation, persisted by a checkpointer so the thread can be resumed, and long-term memory, which is kept in a store and shared across threads. In the terms of this post, short-term memory is the saved context of one conversation, and long-term memory is memory proper: it outlives the thread and has to be looked up to matter.
Deciding what goes where
Ask four questions about any piece of information an agent needs.
- Is it needed on almost every call, and is it short? Put it in always-loaded memory: a rules file, standing instructions.
- Is it needed sometimes, and can you say when? Keep it in memory behind an index or a tool, so it is loaded only then.
- Is it only needed for this task? Leave it in context and let it go. Writing it down just creates something to go stale.
- Do other people, or other agents, need to see it and know who wrote it? Put it in a shared record with history, not in one agent’s private notes.
A useful test: if the information changes every day, it does not belong in a file that is loaded whole. Status, assignments and progress are state, and state belongs somewhere that can be queried for its current value.
Where it usually goes wrong
- Memory written, never read. Notes pile up in a folder nothing opens. Fix the way back in before adding more to the store.
- Memory poured into context. Everything is loaded “just in case”, the window fills with old detail, and the current task gets less attention.
- Stale memory read as fact. A note from last month says a migration is pending; it shipped two weeks ago. Date what you keep, and prune it.
- Private memory for shared work. One agent remembers a decision the rest of the team never saw, and nobody can tell where it came from.
How fenbs splits it
On a fenbs board the two kinds are kept apart on purpose. AI context is the short, always-read part: notes that apply to everything, plus notes for one project that are only given to an assistant working on that project. An assistant reads them with fenbs_get_context at the start of a session, and can add one with fenbs_add_context_note, signed with its name. The tasks themselves are the queried part: the assistant calls fenbs_list_items or fenbs_get_item when it needs them, so a board of hundreds of tasks costs only the few it asks for. Every change to either is recorded with who made it, so what the agent remembers is something you can read too.
Related reading
What is agent memory? covers the kinds of memory in general, and agent memory over MCP covers serving it from a server. For the board side, see AI context and connecting Claude to your board.