MCP vs RAG: Different Jobs, Often Used Together
RAG decides what text the model reads before it answers. MCP gives the model tools it can call while it works. One is about knowledge, the other about reach, and an MCP server is often where the retrieval lives.
7 min read
RAG and MCP answer different questions. Retrieval-augmented generation is a technique: before the model answers, your application searches a body of documents and puts the most relevant passages into the prompt, so the model answers from them rather than from memory. MCP is a protocol: it lets an assistant discover tools on a server and call them while it works, including tools that change things. RAG is about what the model knows at the moment it answers; MCP is about what it can reach and do. They are not alternatives, and a common setup uses both: an MCP server whose search tool is a retrieval pipeline underneath. If you need the protocol explained first, see what MCP is.
The short version
- What it is: RAG is a pattern you build into an application; MCP is a wire protocol between an application and a server.
- Who decides: in classic RAG your code retrieves before the model sees the question; with MCP the model decides mid-task whether to call a tool at all.
- Direction: RAG only reads; MCP tools can read or write, so an assistant can file a task as well as find one.
- Freshness: RAG answers from an index that is as fresh as its last update; an MCP tool usually queries the live system.
- Where they meet: a retrieval pipeline can sit behind an MCP tool, so any MCP client can use it without its own copy of the index.
What RAG actually does
The term comes from a 2020 research paper, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (opens in a new tab), which combined a model’s own trained knowledge with a searchable store of documents. In practice today it means a pipeline you run before each answer.
- Split the source material into chunks, typically a few hundred tokens each.
- Turn each chunk into an embedding and store it in a vector index, often alongside a keyword index.
- When a question arrives, search both indexes and merge the results.
- Put the top few chunks into the prompt and ask the model to answer from them.
Anthropic’s write-up on contextual retrieval (opens in a new tab) walks through that pipeline and a way to improve it: prepend a short explanation of where each chunk sits in its document before indexing it, so a chunk that says “revenue grew 3%” still carries which company and which quarter. The same post makes a point worth remembering before you build anything: if the whole knowledge base is under about 200,000 tokens, roughly 500 pages, you can put all of it in the prompt with prompt caching and skip retrieval altogether.
What RAG does not do is act. It fills the prompt with text and the model reads it. The model does not choose the query, cannot ask a follow-up search when the first one misses, and cannot change anything in the system the text came from.
What MCP adds
MCP moves the decision to the model. A server publishes tools, each with a name, a description and an input schema; the client lists them and the model calls whichever one helps, as often as it needs. The MCP tools specification (opens in a new tab) calls tools model-controlled for exactly this reason, and says there should always be a person able to deny a call, because tools can do more than read.
That changes three things compared with RAG. The model can search, look at what came back, and search again with better words. It can combine search with action: find the open task about the login bug, then comment on it. And it reads the live system rather than a snapshot, so a task moved to Completed five minutes ago shows up as completed.
MCP also has a primitive closer to RAG’s spirit. MCP resources (opens in a new tab) are pieces of context a server offers by URI, such as files or schemas, and they are application-driven: the host decides whether to show them in a picker, let the person choose, or include them automatically. A resource is something to read; a tool is something to call.
Pre-loading versus fetching on demand
The real design choice is not RAG or MCP but when context arrives. Anthropic’s guide to context engineering for agents (opens in a new tab) contrasts embedding-based retrieval before the model runs with a just-in-time approach, where the agent keeps lightweight references such as file paths or stored queries and loads data through tools as it goes. It describes Claude Code as a hybrid: instructions files go in up front, and search tools fetch files when the work needs them. Our own guide to context engineering covers that balance in more depth.
- Pre-load (RAG-style) when the question is predictable and the answer is in a fixed corpus: a help centre, a policy library, product documentation.
- Fetch on demand (MCP tools) when the task is open-ended, the data changes by the minute, or the assistant will need to act on what it finds.
- Do both when a small, stable core always matters and the rest is large or changing: load the core, expose the rest as a tool.
An MCP server can serve retrieval
The two meet cleanly when you wrap a retrieval pipeline in a tool. The model calls search_docs with a query; the server runs the embedding search, reranks, and returns the top passages with their sources. The model gets RAG’s results, but on its own schedule, and every MCP client can use the same index without each one building it again.
tools/call search_docs { "query": "refund window for annual plans", "limit": 5 }
-> server: embed query, search vector + keyword index, rerank
<- content: five passages, each with its document title and URIA few things to get right when you do. Return the source with each passage so the assistant can cite it. Keep results short, since every passage is context the model has to read. Say in the tool description what the index covers and how fresh it is, so the model knows when to trust it and when to look elsewhere. And apply the caller’s permissions at query time: the specification lets a server vary its answers by the authorization on each request, and a retrieval tool that ignores who is asking will happily return documents the person could never open themselves.
Not all retrieval needs embeddings
Many useful search tools are plain keyword search over structured records, and for work items that is often the better fit: people search for a task by the words in its title or a reference like BUG-142, not by meaning. On fenbs, fenbs_search searches tasks across every board the caller can see. It matches the title, the note, the plan, the task reference, and project and category names; with several words, every word must appear somewhere, in any of its simple forms. There is no vector index behind it. It is a live query of the board, so it can never be out of date, and it only returns what the caller is allowed to see.
fenbs also does something RAG-shaped in the ordinary sense. fenbs_get_context returns the board’s AI context: the standing notes people and earlier assistants left for every assistant to read before it starts. An assistant calls it once at the start of a session, which is the pre-load; it then searches and lists tasks as the work needs them, which is fetching on demand. And fenbs_create_item runs a similarity check against open and recently finished tasks before filing, so a search the assistant forgot to do still happens.
How to choose
- A chatbot answering questions from a fixed document set, with no actions: RAG, or the whole corpus in the prompt if it is small enough.
- An assistant that works inside a live system, reads current records and changes them: MCP tools.
- An assistant that needs a large document set while it works: a retrieval pipeline exposed as an MCP tool.
- A small amount of guidance that applies to every session: load it up front, whether from an instructions file or a context tool.
Related
The other comparisons in this series: MCP vs API, MCP vs function calling and MCP vs Claude Skills. Building a server with a search tool of your own: how to build an MCP server. The full list of fenbs tools: MCP docs.