What Is Supermemory? Memory and Context Infrastructure for AI Agents
Supermemory is a memory and retrieval service for AI agents: you send it documents and conversations, and it builds a graph of facts that update over time, user profiles and searchable chunks behind one API. How it works, how it isolates users, what runs over MCP, and who it is for.
7 min read
Supermemory is context infrastructure for AI agents. It gives an agent long-term memory, retrieval over documents and automatically maintained user profiles through one API, so a developer sends it raw material and gets back what the agent should know. You hand it documents, which can be chat transcripts, PDFs, web pages or files synced from a user’s Google Drive or Gmail, and it extracts facts that link to each other and change as new information arrives. It is mainly a developer product, with a hosted platform, SDKs for TypeScript and Python, a self-hosted binary and a hosted MCP server. Everything below is from Supermemory’s own documentation and changelog as of October 8, 2026.
For the general idea of agent memory, start with what agent memory is. This page covers one product.
Documents in, memories out
Supermemory separates what you send from what it keeps. A document is raw input. Its guide to how Supermemory works (opens in a new tab) lists the pipeline each document passes through: queued, extracting text or transcribing, chunking, embedding, indexing, done. You do not pre-chunk anything or choose an embedding model. The chunks stay searchable, which Supermemory sells as managed retrieval-augmented generation under the name SuperRAG, and the engine also extracts memories: short facts about a user or an entity.
import { Supermemory } from "supermemory";
const supermemory = new Supermemory({ apiKey: "sm_..." });
await supermemory.add("user_123", {
content: "The user loves Paris.",
});
const { results } = await supermemory.search("user_123", {
query: "where does the user want to travel?",
});The first argument, user_123, is the namespace: the boundary everything in that call belongs to.
A memory graph that changes over time
Supermemory’s documentation on graph memory (opens in a new tab) describes memories that connect in three ways as content arrives:
- Updates: “Alex just started at Stripe as a PM” replaces “Alex works at Google” for search, while the history can remain.
- Extends: “Alex leads a team of five” adds detail without invalidating anything.
- Derives: from several memories, the engine infers a fact nobody stated in one place.
Inferred facts are flagged and down-weighted in search until someone approves them, and an API lists them so an application can build a review screen: approve, decline, or undo a decision. Memories can also be forgotten, exactly or by meaning. User profiles sit on top: stable and recent facts about a person, fetched in one call rather than searched for.
Namespaces: how users are kept apart
Every call is scoped to one namespace (opens in a new tab), called a container tag in older versions of the API. A namespace is an identifier you choose, such as user_alex or org:acme:user:john, and the first write creates it. Each one is backed by its own vector index, so one user’s memories are never searched alongside another’s, and a request can touch only one namespace. Deleting a user is a matter of deleting their namespace and revoking any key scoped to it.
Connectors and integrations
On the hosted platform, connectors sync content in the background from Google Drive, Gmail, Notion, OneDrive, GitHub, Granola, Amazon S3 and a web crawler. For developers there are integration guides for the Vercel AI SDK, LangChain, LangGraph, the OpenAI Agents SDK, Mastra, CrewAI and others, a backend for Claude’s memory tool, and plugins for coding tools such as Claude Code and Codex.
Supermemory over MCP
The Supermemory MCP server (opens in a new tab) at https://mcp.supermemory.ai/mcp lets any MCP-compatible assistant use the same memory. You sign in with OAuth, so there is no API key to paste, and choose which spaces the client can access. Spaces keep a team’s material for one initiative together, and teammates work within the spaces they may read or write. Its tools include search_memory, add_memory, which saves or forgets, get_profile, list_documents and list_memories. If MCP is new to you, see what MCP is.
Self-hosting, security and the consumer side
Supermemory local is a single binary that runs the same API on your machine, with built-in embeddings and any OpenAI-compatible model, offline if you like. Its documentation is direct about the limits: the SDKs are open source, but the server binary is built from a non-public codebase, is free within a lite license limit, and leaves out connectors and the Supermemory MCP.
For the hosted service, Supermemory’s security and compliance page (opens in a new tab) lists SOC 2 Type II, GDPR, a HIPAA business associate agreement on eligible plans, encryption in transit and at rest, and says customer content is never used to train models on any plan.
There is a personal side too. Supermemory’s June 24, 2026 changelog (opens in a new tab) says its Chrome extension guides imports of memory from ChatGPT, Claude, Grok and Gemini. For most readers, though, Supermemory is something a developer builds on.
Getting started
- Create an API key in the developer console, under API Keys.
- Install the SDK:
npm install supermemoryorpip install supermemory. - Choose your namespace convention, such as one per user or one per customer workspace, before the first write.
- Add a conversation or a document, then search it with a question your agent would really ask.
- Decide where wrong or inferred memories will be reviewed, by your team or by your users, before you go live.
To try it without the hosted service, npx supermemory local starts the self-hosted server, which prints an API key on first boot; point the SDK at it by changing its base URL.
Who should use Supermemory
- Teams building an AI product whose agent must remember each user and also search that user’s files and inbox.
- Multi-tenant apps where strict separation between customers is a hard requirement.
- Developers who want retrieval and memory from one service rather than a vector database plus glue code.
If your memory comes only from conversations, or you need an open-source server, compare it with Mem0 first; Mem0 vs Supermemory goes through the differences.
Two other cases point elsewhere. If you want an agent that keeps and edits its own memory as files you can read, Letta is built around that idea. If policy over company data, who may retrieve what and where each fact came from, is the hard part, Zep is built around governance. Letta vs Mem0 vs Zep compares those designs.
What it is not
Supermemory decides what to remember: a learning model extracts, links and forgets. That is what you want inside a product used by thousands of people, and it is fast, scalable and well isolated. It is a different thing from a record that a team writes deliberately and reads directly: the rules for how work is done, the decisions behind them, the tasks in progress and where the last session stopped.
fenbs is a web app for that record. Your projects, tasks, Decisions and rules, lessons learned and Where we left off sit on one board people can see and edit, and ChatGPT, Claude, Claude Code, Codex, Cursor and other AI apps read and update it over MCP. People write the rules and every assistant reads them first; assistants can add notes, signed with their name, but cannot decide a decision or pre-approve work, and History records who changed what. If you are building an AI app for your users, look at Supermemory. If you use AI for your own work across several apps, keep the shared record on a board.
Related
Other engines: what is Mem0, Mem0 alternatives and Letta vs Mem0 vs Zep. The risks of stored memory: agent memory security. Notes every assistant reads on fenbs: AI context.