Choosing an Agent Memory System: Criteria That Matter
There is no best agent memory system in general, only the one that fits who reads the memory, who writes it and who has to fix it when it is wrong. Seven criteria, and how six common approaches score against them.
8 min read
The best agent memory system is the one that scores well on the criteria your situation cannot do without, and those depend less on technology than on people. Judge any option on seven questions: what it stores, who can read it, who can write and correct it, how well the right memory comes back at the right moment, whether it moves with you between assistants, whether you can see who changed what, and where private data ends up. Then add the cost of keeping it running. Instruction files, Claude Code’s auto memory, Anthropic’s memory tool, LangGraph stores, vector databases, MCP memory servers and a shared board each win on some of these and lose on others. This guide scores them structurally, from their own documentation, so you can pick with your eyes open.
If you are still deciding what memory is for, start with what agent memory is. If you have already chosen and want to build, how to build agent memory is the step-by-step. This page sits between the two: how to choose.
Start with the situation, not the product
Three situations cover most teams, and each puts a different criterion first:
- One person, one assistant, one machine. Simplicity and ease of correction come first. Anything you cannot open and edit in a minute is overkill.
- An application serving many users. Isolation comes first: one user’s memories must never reach another. Retrieval quality and running cost follow.
- A team with several people and several assistants. Shared reading, attribution and portability come first, because the memory has to outlive any one assistant and any one laptop.
The seven criteria
- What is stored. Rules you wrote, notes the agent wrote, whole documents, or all three. A system built for documents is a poor place for a one-line correction, and the reverse.
- Who can read it. One session, one machine, one user, or everyone on a team. Read access is also what decides where a leak can go.
- Who can write and correct it. The agent, a person, or both, and how easily a wrong entry is fixed or removed. A memory that is hard to correct keeps being wrong.
- Retrieval quality. Is the memory read whole, fetched by name, or found by similarity search? Read whole is exact but must stay small; search scales but can miss the entry that mattered.
- Portability. Does it work with only one client, or can Claude, Cursor and whatever you use next year read the same memory?
- Audit. Can you tell who wrote an entry, when, and from which session? Without that, a wrong memory looks exactly like a right one.
- Privacy and running cost. Where the data sits, who operates the store, and what it takes to keep it pruned, backed up and secure.
How six approaches score
Instruction files: CLAUDE.md and AGENTS.md
Files you write and commit. According to Anthropic’s documentation on how Claude remembers your project (opens in a new tab), CLAUDE.md files can be scoped to an organisation, a user or a project, are loaded at the start of every session, and are treated as context, not enforced configuration. The same page says Claude Code can now read a repository’s AGENTS.md as its project instructions when there is no CLAUDE.md.
Scores: excellent on correction and audit, since every change is a commit with an author. Good on portability, because AGENTS.md (opens in a new tab) is an open format read by many coding agents. Weak on who writes it: the agent can propose edits, but the file is really yours. Retrieval is read-whole, so it must stay short.
Claude Code auto memory
Notes Claude writes itself, from your corrections and preferences, in a folder per repository under ~/.claude/projects/. The documentation says the first 200 lines or 25KB of its MEMORY.md index load at the start of every conversation, topic files are read on demand, and the memory is machine-local: it is not shared across machines or cloud environments. You can toggle it from /memory.
Scores: strong on effort, since nobody has to write anything, and good on correction, since the files are plain Markdown you can edit. Weak on portability and sharing by design: it serves one person, one client, one machine.
Anthropic’s memory tool
A tool in the Claude API rather than a store. The memory tool documentation (opens in a new tab) says it runs client-side: Claude requests operations such as view, create, str_replace and delete on a /memories path, and your application carries them out against storage you choose, such as a per-user directory or keys in a database. It also says the safeguards are yours: stripping sensitive data, capping file size, expiring old files and blocking path traversal.
Scores: whatever you build. Isolation, audit and privacy are as good as your handler. That is a strength for a product team and a cost for anyone who wanted memory without writing a backend.
LangGraph stores
LangGraph’s guide to adding memory (opens in a new tab) separates short-term memory, kept per conversation thread by a checkpointer, from long-term memory in a store. Long-term items are put under a namespace, such as a user id, and fetched by key or searched, optionally by embedding similarity. It recommends an in-memory store for development and lists database-backed stores such as Postgres and Redis for production.
Scores: strong on isolation, because the namespace is the boundary between users, and on retrieval choice. Portability is to your application, not across assistants: a person using Cursor next door does not see it.
Vector databases
The usual engine behind retrieval: text is split into chunks, embedded and searched by similarity. It scales furthest and is the hardest for a person to inspect or correct, since fixing a memory means finding and replacing its source and re-indexing. Multi-user setups carry a specific risk: OWASP’s entry on vector and embedding weaknesses (opens in a new tab) warns of context leaking between users who share one vector database, and recommends permission-aware stores with strict partitioning. When memory and retrieval are the right tool for which job is covered in agent memory vs RAG.
MCP memory servers
A server that exposes memory as tools any MCP client can call. The reference knowledge-graph memory server (opens in a new tab) stores entities, relations and observations in a local JSONL file and offers tools to create, delete, read and search them. It has no sign-in and no per-user separation; a remote server you build can add both. Scores: good on portability between clients, and as good on audit and access as the server behind it. The options are compared in agent memory over MCP.
A shared board
Memory kept in the system the team already works in, with members, roles and a history. On fenbs this is AI context: short notes that every connected assistant reads with fenbs_get_context, either for everything or for one project, where a project’s notes are only given to an assistant working on that project. An assistant adds a note with fenbs_add_context_note, signed with its name, and corrects one with fenbs_update_context_note; changes are recorded in History. Anyone who can see the board can read the notes, and changing them needs the permission to change tasks.
Scores: strong on sharing, correction, audit and portability, because Claude, ChatGPT, Cursor and the rest read the same notes over MCP. Weak on volume by design: these are short notes meant to be read whole, not a document store or a vector index, and a deleted note cannot yet be restored from fenbs.
The scorecard
Share Correct Audit Portable Isolation Volume Instruction files repo easy commits wide per repo small Auto memory you easy files one tool per machine small Memory tool (API) you build it: all six depend on your handler LangGraph store app code yours one app namespace large Vector database app hard yours one app partition largest MCP memory server server varies varies any MCP varies medium Shared board team easy history any MCP project small
Read it by row against your situation. One person on one machine is well served by the first two rows. A product for many users needs one of the middle rows, with isolation built before anything else. A team needs at least one row where the memory is shared and every change has a name.
Questions to ask of any memory system
- Show me one memory entry. Can I see who wrote it and when?
- How do I correct a wrong entry, and how long until every assistant stops using the old one?
- What stops user A’s memory reaching user B?
- Which assistants can read it today without custom code?
- Where is the data stored, and who can read it outside the assistant?
- How does an entry expire, and who is responsible for pruning?
An option that cannot answer the first two is fine for personal preferences and risky for anything a team relies on. The security side of those answers, including what should never be stored, is in agent memory security.
Combinations that work
Most setups use two or three layers rather than one system. A common shape: instruction files for rules that belong to the code, personal auto memory for one person’s preferences, a shared record for what the whole team and every assistant must know, and a retrieval store only for reference material too large to read whole. Each layer then has one job, one owner and one place to correct it.
Related
The kinds of memory: what is agent memory. Building your own: how to build agent memory. Memory in Claude Code specifically: giving Claude Code a memory. Connecting an assistant to a board: the MCP docs.