Agent Memory vs RAG: What Each Stores, Who Writes It, When It Is Read
RAG fetches passages from documents that existed before the agent arrived. Agent memory holds what was learned while working. They fail in different ways, need different owners, and most long-running agents need both.
7 min read
Agent memory and RAG both put knowledge into a model’s context window that was not there before, but they hold different things. Retrieval-augmented generation fetches passages from a body of documents that existed before the agent arrived: manuals, policies, tickets, a codebase. People wrote them, a pipeline indexed them, and a search picks the most relevant pieces for each question. Agent memory holds what was learned while working: a preference, a correction, a decision, how far a job has got. The agent or the people directing it wrote it, it is small enough to read at the start of a session, and it is changed in place when it stops being true. RAG is the library; memory is the notebook. For the kinds of memory in general see what agent memory is, and for how retrieval relates to tools, MCP vs RAG.
The difference at a glance
- What it holds. RAG: chunks of source documents. Memory: short, distilled facts, decisions, preferences and progress.
- Who writes it. RAG: the authors of the documents, then an indexing pipeline. Memory: the agent itself during the work, or the people it works with.
- When it is read. RAG: on every question, by similarity search. Memory: a small index at the start of each session, with detail read when needed.
- How it changes. RAG: fix the source document and re-index. Memory: edit or delete the note, often done by the agent.
- How it fails. RAG: the passage that mattered is not fetched, or a near miss is. Memory: a stale or wrong note is acted on as if it were true.
- Size. RAG: as large as the corpus. Memory: small by design, because it is read whole or nearly whole.
RAG was always called memory
The confusion is old. The 2020 paper that named the technique, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (opens in a new tab), describes its models as combining “parametric” memory, the knowledge in the model’s weights, with “non-parametric” memory, a dense vector index of Wikipedia searched by a retriever. So RAG is a kind of memory. What separates it from agent memory is not the storage technology but the direction of writing: in RAG, the agent only reads.
Anthropic’s write-up on contextual retrieval (opens in a new tab) describes the usual pipeline today: split the material into chunks, create embeddings for semantic search and a BM25 index for exact words and phrases, fetch the best matches for each question and add them to the prompt. It also makes a point worth checking before building anything: if the knowledge base is under 200,000 tokens, about 500 pages, you can put all of it in the prompt and skip retrieval.
- Good for: large reference material that people maintain, where the answer is written down somewhere.
- Weak at: anything nobody wrote down, such as what this user prefers or what was tried yesterday, and questions whose wording is far from the words in the answer.
What agent memory is for
Agent memory is for what the documents do not say. Anthropic’s memory tool (opens in a new tab) for the Claude API shows the shape: with it enabled, Claude checks its memory directory before starting a task, stores what it learns in files under /memories as it works, and reads them back in later conversations. The files are created, edited and deleted by the model through tool calls that your application carries out.
Claude Code’s auto memory makes the boundary explicit. According to Claude Code’s memory documentation (opens in a new tab), Claude saves notes about you, your corrections, ongoing work and where to find things, and skips anything it can derive from the codebase. In other words, memory stores what retrieval over the code could never find, because it is not in the code.
- Good for: preferences, corrections, decisions, the state of a job, facts learned by trial and error.
- Weak at: volume. Memory that is read whole must stay short, and memory nobody prunes grows stale.
Who writes it changes who you trust
In RAG, the sources were written and usually reviewed by people, and the agent cannot change them. A wrong answer traces back to a document with an owner, and the fix is to correct the document and re-index it. That is slow, but the fix lands in the right place.
With memory, the agent writes to the store it will later read. A wrong inference saved once is read as a fact in every later session, with no reviewer in between. That is why memory needs things RAG can do without: a visible date on each note, a named author, pruning, and an easy way for a person to correct or delete an entry. The memory tool documentation lists some of the safeguards as your responsibility: strip sensitive data before a file is written, cap how large files grow, and periodically delete files that have not been accessed in a long time.
When memory turns into retrieval
Memory stores grow, and a large one runs into the same problem RAG solves. Claude Code handles it by loading only the first 200 lines or 25KB of its MEMORY.md index at the start of a session and reading the topic files it points to on demand. That is retrieval, by file name rather than by embedding. Anthropic’s guide to effective context engineering (opens in a new tab) describes the same pattern as structured note-taking: the agent keeps notes outside the context window and pulls them back in when they are needed, alongside “just in time” strategies that keep lightweight references such as file paths and load the data through tools.
Once memory is searched rather than read, it inherits retrieval’s failure: the note that mattered may not be fetched. The better fix is usually to keep memory small and pruned, so it can still be read whole, and to move reference material that has outgrown it into documents that a retrieval pipeline can index.
Using both
Most agents that run for longer than one session need both, for different things. A support assistant retrieves from the help centre and remembers that this customer was promised a call back. A coding agent searches the codebase with grep as it works and remembers that the integration tests need a local database running. A simple rule decides where each fact goes:
- It existed before the agent, and people maintain it: retrieval.
- The agent learned it, or someone told it: memory.
- Other people or other assistants need it: a shared record, not one agent’s private memory.
- It changes by the minute, like the status of a task: neither. Query the live system through a tool.
The third rule is the one most setups miss. Memory files and memory tools usually belong to one assistant on one machine. When the agent’s notes matter to a team, they belong somewhere the team can read and correct them.
Where a board fits
A fenbs board covers the shared-memory and live-state parts. Its AI context is memory that every connected assistant reads with fenbs_get_context before it starts: short notes from people and earlier assistants, for everything or for one project, each signed with who wrote it and changed in place with fenbs_update_context_note when it goes out of date. The state of each piece of work lives on its task, in the note, the plan and the comments. Finding a task is a live keyword search with fenbs_search, not a vector index, so it cannot go stale. A board is not a document store, so large reference material still belongs in a retrieval pipeline or a documentation server of its own.
Related
The wider practice of choosing what goes into the window: context engineering for AI agents. Memory served by an MCP server: agent memory over MCP. Connecting an assistant to a board: the MCP docs.