Agent Memory: Short-Term, Long-Term and Shared

An AI agent starts every session with an empty context window. Agent memory is whatever puts the right knowledge back. The main kinds, what each is good for, and who can see and change what the agent remembers.

7 min read

Agent memory is anything that lets an AI agent use knowledge beyond what it was trained on and beyond the current moment: what it has seen earlier in this session, and what it, you or someone else wrote down in an earlier one. There are four common kinds. Short-term memory is the context window itself. Long-term memory is usually files, such as instructions you write and notes the agent writes for itself. Retrieval memory is a searchable store, often a vector database, that the agent queries for relevant passages. Shared memory is a record that several agents and people read and write, such as a repository or a task board. They differ less in technology than in who writes them, who can read them, and who can correct them when they are wrong.

Why agents need memory at all

A language model has no memory between requests. Each call is answered from what is in its context window at that moment, and a new session begins with that window empty. Anthropic’s write-up on harnesses for long-running agents (opens in a new tab) puts the problem in one line: each new session begins with no memory of what came before. For a job that fits in one sitting, that is fine. For a project that runs over days, or passes between assistants and people, something has to carry the knowledge across.

Short-term: the context window

Everything in the current window is, in effect, the agent’s short-term memory: your instructions, the files it has read, tool results and the conversation so far. It is the only memory the model uses directly; every other kind works by putting something back into this one.

  • Strengths: immediate, exact and needs no set-up. If it is in the window, the model can use it.
  • Weaknesses: limited in size, lost at the end of the session, and less reliable as it fills. When a long session is compacted, what was said only in conversation comes back as a summary.
  • Who can see and edit it: whoever is in the session. Nobody else can read it, and you correct it by saying so, which adds more to the window rather than removing the mistake.

Long-term: memory in files

The simplest durable memory is a file the agent reads at the start of each session. Claude Code has two built-in versions, described in its documentation on how Claude remembers your project (opens in a new tab):

  • CLAUDE.md files, written by you. They can be scoped to an organisation, to you across all projects, or to one project, and a project file is shared with the team through source control. Claude treats them as context, not as enforced rules.
  • Auto memory, written by Claude from your corrections and preferences. It is kept per repository in a folder under ~/.claude/projects/, and it is machine-local: it is not shared across machines or cloud environments. The first 200 lines or 25KB of its MEMORY.md index load at the start of each session.

Developers building their own agents on the Claude API can use the memory tool (opens in a new tab), which gives the model file operations on a /memories directory. It runs client-side: Claude asks for a read or a write and your application carries it out against storage you control, so where the memory lives, how long it is kept and who else can reach it are your decisions.

File memory is easy to understand and easy to fix. You can open the file, read it and delete a wrong line. Its weaknesses are size, since it all loads every time, and scope, since a file on one laptop helps nobody else. How to lay out instruction files across several repositories is covered in managing multiple projects with CLAUDE.md.

Retrieval: vector stores and search

When there is too much to load every time, such as a documentation site, years of support tickets or a large codebase, the usual answer is retrieval. Anthropic’s post on contextual retrieval (opens in a new tab) describes the standard pipeline: split the material into chunks, turn each chunk into an embedding that encodes its meaning, store the embeddings in a vector database, and at run time fetch the chunks most similar to the question and add them to the prompt. The same post notes that a knowledge base under about 200,000 tokens can simply go into the prompt whole.

  • Strengths: scales to far more material than any window, and fetches only what looks relevant.
  • Weaknesses: “looks relevant” is a similarity score, not understanding, so it can miss the passage that mattered or fetch a near miss. It is also hard to see what the agent will recall for a given question.
  • Who can see and edit it: whoever runs the pipeline. Correcting a wrong memory means finding and replacing the source document and re-indexing it, which is rarely something the person using the agent can do.

Agents also retrieve without vectors. A coding agent that runs grep or reads a file when it needs it is doing just-in-time retrieval over the filesystem, which Anthropic describes in its guide to context engineering (opens in a new tab) alongside structured note-taking: the agent writing notes to memory outside its window and reading them back later.

Shared: records that outlive the agent

The last kind is not memory built for an agent at all. It is the record the team already keeps: the repository and its history, a progress file, the tasks on a board, the decisions log. Anthropic’s long-running agent harness leans on exactly this, having each new session read the git log and a progress file before it does anything. These records have a property the others lack: they belong to the work, not to one assistant, so a job started in Claude Code can be picked up in Cursor or by a colleague.

  • Strengths: shared by default, readable by people as well as agents, and each change has an author and a date.
  • Weaknesses: someone has to keep them current, and an agent only benefits if it is told to read them.
  • Who can see and edit it: whoever has access to the record, under the same permissions as everyone else.

Who can see and edit each kind

The question that decides which memory to use for what is rarely technical. It is about audience and correction:

  • The context window: one person, one session. Gone when the session ends.
  • Auto memory and personal files: one person on one machine. Readable and editable, but invisible to the rest of the team.
  • Project instructions files: everyone with the repository, changed through commits and reviewed like code.
  • Vector stores: readable by the agent, maintained by whoever built the pipeline. Hard for users to inspect or correct.
  • Task boards and shared records: everyone with access, each change attributed, correctable by anyone permitted to edit.

Two cautions apply to all of them. Memory goes stale, and an agent will act on an out-of-date note as confidently as a current one, so every memory needs an owner who prunes it. And memory can leak: the memory tool documentation advises validating what gets written so that sensitive information is stripped out. Secrets and personal data do not belong in anything an agent reads back.

Shared memory on a fenbs board

A fenbs board is the shared kind. Standing knowledge lives in the board’s AI context: short notes, for everything or for one project, that each connected assistant reads when it starts. An assistant that works something out can leave a note for the next one, signed with its name; if a note is out of date it should be changed rather than contradicted. Anyone who can see the board can read the context, and changing it needs the permission to change tasks.

Reading and adding to the board’s shared memory over MCP
fenbs_get_context         project: "api"      # notes for everything, plus this project
fenbs_list_context_notes  project: "api"      # find a note to change
fenbs_update_context_note id: 12, body: "…"   # correct it instead of adding a contradiction
fenbs_add_context_note    title: "Tests need the API running", body: "…"

The state of each piece of work sits on its task: the note (what the problem is), the plan (how it will be done, rewritten as the agent learns) and the comment thread. Changes are recorded in the board’s history with who made them, so when an assistant remembers something wrongly you can see where it came from. The wider practice of deciding what goes into the window, and when, is in context engineering for AI agents.

Related

For how long sessions degrade and why writing state down helps, see context rot. For a day-to-day routine of reading and updating a board, see a task-tracking workflow for Claude Code. To connect an assistant, start with the MCP docs.

Questions people ask.

What is agent memory in simple terms?

It is whatever lets an AI agent use knowledge from outside the current moment: what is in its context window now, and notes, files, search indexes or shared records that put earlier knowledge back into the window when it is needed.

What is the difference between short-term and long-term agent memory?

Short-term memory is the context window of the current session, and it is lost when the session ends. Long-term memory is stored outside the model, usually in files or a database, and loaded or searched at the start of later sessions.

Is Claude Code auto memory shared with my team?

No. According to Claude Code’s documentation, auto memory is kept per repository on your machine and is not shared across machines or cloud environments. Project CLAUDE.md files, by contrast, are shared with the team through source control.

Which kind of memory should a team rely on?

For anything more than one person needs, such as rules, decisions and the progress of work, use a shared record everyone can read and correct, and keep personal memory for personal preferences. Retrieval stores suit large reference material that is too big to load whole.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.