How to Build Memory for an AI Agent
Five decisions make up an agent’s memory: what it keeps, where it keeps it, who writes it, how it gets back into context, and how it is corrected. A build guide, one decision at a time, with the failure each choice invites.
8 min read
To build memory for an AI agent, make five decisions in order. Decide what is worth keeping, and write down what is not. Pick a store that matches who needs to read it: files for one agent, a database for many users, a shared board for a team, a vector store only when there is too much to read whole. Decide who writes each memory, and whether a person reviews it before it counts. Decide how memory gets back into the context window, because a note nobody loads changes nothing. And decide how a memory is corrected and when it expires. Most memory that disappoints was built storage-first and skipped the last two. If you want the kinds of memory explained before you build, start with what agent memory is; this guide assumes you have decided you need some.
Decision one: what to remember
Start with a list, not a database. Write down the things your agent got wrong last week because it did not know them. Those are your first memories. They usually fall into four groups:
- Preferences and corrections: “use pnpm, not npm”, “this client wants replies under 150 words”.
- Facts about the environment that are not in the material the agent reads: “staging resets nightly”, “the payments sandbox needs a VPN”.
- Decisions, with who made them and why, so the agent does not undo something deliberate.
- Progress: what was finished, what was tried and failed, what comes next.
LangChain’s memory concepts page (opens in a new tab) offers a useful vocabulary for the same split: semantic memory holds facts, episodic memory recalls past events or actions, and procedural memory holds the rules used to perform tasks. It also distinguishes a profile, one continuously updated record about a user or entity, from a collection of many small documents. A profile is easier to keep true; a collection is easier to add to and harder to correct. For most agents, start with a profile per user or project and a short collection of decisions.
Just as important is the list of what never goes in. Anything the agent can derive by reading the code or the documents: storing it twice means two versions that will drift. Anything that changes by the minute, such as the status of a task: query it live instead. And secrets, tokens and personal data, always.
Decision two: where memory lives
Choose the store by audience first and technology second. Four options cover almost every case:
- Files. A folder of Markdown notes the agent reads and edits. Easy to inspect, easy to diff, easy to delete a wrong line. Right for one agent and one person.
- A database or key-value store. One record per user, project or tenant, keyed so each only ever sees its own. Right when you run the agent for many users and need isolation you can prove.
- A task board or other shared record. Right when people and several assistants need the same memory, and each change needs an author.
- A vector store. Right only when there is more material than you could ever load and no obvious index. It retrieves by resemblance, so it is better for reference material than for facts that must be current.
The vector store is the one most often chosen too early. Memory that is meant to be read whole should be small enough to read whole; if yours is not, prune it before you index it. The trade-off is covered in depth in agent memory vs RAG, and serving any of these from a server that several clients share is covered in agent memory over MCP.
Decision three: who writes it
There are three writers: the agent, a person, or the agent with a person approving. Each changes what you can trust.
If the agent writes, decide when. The same LangChain page describes two timings: writing in the hot path, where a new memory is available at once but the agent slows down while it decides what to save, and writing in the background, where a separate process distils memories after the conversation and you choose how often it runs. Hot-path writing suits corrections the user expects to stick immediately. Background writing suits summaries of long sessions.
If a person reviews, make review cheap. The pattern that works is proposal then acceptance: the agent writes a short proposed note, a person reads it in the same place they read everything else, and only accepted notes are loaded next time. Keep each note in a shape a person can judge in five seconds:
--- written_by: assistant (for Priya) written_at: 2026-09-28 source: session on the checkout retry bug status: proposed # proposed | accepted | superseded --- Integration tests need the local Redis running (docker compose up redis). Without it they fail with ECONNREFUSED, not with a test error.
Author, date, source and status are what turn a note from something the agent believes into something you can check. Without them, a wrong memory is indistinguishable from a right one.
Decision four: how memory gets back into context
The model only ever reasons over what is in its context window, so retrieval is the half of memory that does the work. The difference between the two is the subject of agent memory vs context; here is how to build the way back in. There are three patterns, and most agents use two of them:
- Always load a short index. A few hundred lines at most, read at the start of every session, each line pointing to a longer note.
- Read detail on demand. The agent opens the note the index points to only when the task needs it.
- Search when there is too much to index by hand. Keyword search first, embeddings when wording varies too much for keywords.
Claude Code is a worked example of the first two. Its documentation on how Claude remembers your project (opens in a new tab) says the first 200 lines or 25KB of the auto memory index, MEMORY.md, load at the start of every conversation, and detailed notes move into separate topic files that are read when needed. Content beyond that threshold is simply not loaded, which is the right failure: an index that grows too long gets shorter, rather than quietly eating the window. How that works day to day is in giving Claude Code a memory.
If you are building on the Claude API, the memory tool (opens in a new tab) gives you the mechanics without inventing a protocol. It is client-side: Claude asks to view, create, edit, rename or delete files under /memories, and your handler carries each request out against storage you choose, such as a per-user directory or keys in a database. With the tool enabled, Claude checks its memory directory before starting a task.
import anthropic
from anthropic.tools import BetaLocalFilesystemMemoryTool
client = anthropic.Anthropic()
memory = BetaLocalFilesystemMemoryTool(base_path="./memory")
runner = client.beta.messages.tool_runner(
model="claude-opus-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Remember that Acme prefers email follow-ups."}],
tools=[memory],
)
print(runner.until_done().content)For many users, scope the store before anything else. LangGraph’s guide to adding memory (opens in a new tab) organises long-term memories under a namespace such as a user id plus “memories”, puts new ones with a key, and searches within the namespace, optionally by embedding similarity. Whatever you use, the namespace is your isolation boundary: an agent serving one user must not be able to name another’s.
Decision five: expiry and correction
An agent acts on an out-of-date note as confidently as a current one, so memory needs a way to die. Build three things in from the start:
- A date on every note. Claude Code stamps a
modifiedtime into the frontmatter of memory files it writes, so both you and Claude can see how old a fact is. Copy the idea. - Correction in place. When a fact changes, rewrite the note and mark the old one superseded. Appending “actually, ignore the line above” leaves the agent to pick a winner.
- Expiry. The memory tool documentation suggests periodically deleting memory files that have not been accessed in a long time, and capping how large a file can grow.
Pair memory with compaction for long sessions. Anthropic’s compaction documentation (opens in a new tab) describes summarising older conversation context on the server as a conversation approaches the context limit; the memory tool page recommends using both, since compaction keeps the active context small and memory preserves what must survive the summary. Anything the agent needs after a summary should already be written down.
Guard the write path
Memory is context in waiting: whatever one session writes, every later session reads before it acts. That makes the write path the part to secure.
- Validate paths and keys. The memory tool documentation warns that a path such as
/memories/../../secrets.envcan reach files outside the store unless every path is checked. - Strip sensitive data before a note is stored, not after it has been read back into an answer.
- Treat content the agent read as untrusted. A web page or a ticket can contain text written to be remembered. Reviewed writes, or at least visible authors, are how you catch it.
- Keep an audit trail of memory changes, so a bad note can be traced to the session that wrote it. The wider list is in MCP security risks.
The shared layer: a board as team memory
Files and per-user stores answer “what does this agent know?”. A team also needs “what do we all know, and who said so?”, readable by people and by every assistant they use. On fenbs that layer already exists. AI context holds short notes, for everything or for one project, that every connected assistant reads with fenbs_get_context before it starts. An assistant adds one with fenbs_add_context_note, signed with its name, and corrects one with fenbs_update_context_note rather than adding a contradiction; changes are recorded in History under whoever made them. Decisions have their own page, with who decided and why, and an assistant can write one down but never be the decider.
fenbs_get_context project: "api" # load shared memory fenbs_list_context_notes project: "api" # before adding, look for a note to update fenbs_update_context_note id: 12, body: "Staging resets nightly at 02:00 UTC."
Access follows the same rules as everything else on the board: an assistant signs in as a person, holds that person’s role narrowed by the scopes they ticked, and loses access when the token is revoked. A board is not a document store or an embedding index, so keep large reference material in a retrieval pipeline of its own.
Related
The kinds of memory an agent can have: what is agent memory. Choosing what goes into the window at all: context engineering for AI agents. Connecting an assistant to a board: the MCP docs.