What Is RAG (Retrieval-Augmented Generation)?
RAG is how an AI assistant answers from your documents instead of from memory: find the relevant passages first, put them in front of the model, and ask it to answer from them. How retrieval, chunking, embeddings and citations fit together, with a worked example and a checklist.
8 min read
RAG, short for retrieval-augmented generation, is a way to make an AI model answer from your own documents rather than from what it happened to learn in training. When a question arrives, software first searches a collection of documents for the passages most likely to contain the answer. It then puts those passages into the prompt next to the question and asks the model to answer from them, ideally citing which passage said what. Retrieval finds; generation writes. The model itself is unchanged, which is why RAG is the usual way to give an assistant knowledge of a handbook, a help center or a contract library that changes every week.
If you want to know what the model underneath is and why it needs help with facts, start with what is an LLM. The comparisons have their own posts, linked once in the section on choosing below.
Why a large language model needs retrieval
An LLM writes the most plausible continuation of the text in front of it. It knows nothing about your company, and what it knows about the world stops when its training data does. Ask it about your return policy and it will write a plausible return policy, which is the problem. Google’s overview of its RAG Engine (opens in a new tab), part of what used to be Vertex AI and is now Gemini Enterprise Agent Platform, puts it plainly: a common problem with LLMs is that they do not understand your organization’s private data, and adding that information to the model’s context helps it reduce hallucinations and answer more accurately. RAG is that adding, done automatically for each question.
How RAG works, step by step
The steps are the same whichever vendor you use. Google lists them in this order, and OpenAI’s and Anthropic’s documentation describe the same pipeline in their own words.
- Ingest. Collect the documents: PDFs, help articles, a shared drive, exported tickets.
- Chunk. Split each document into passages small enough that one passage covers one idea.
- Embed. Turn each chunk into an embedding, a list of numbers that captures its meaning, so that passages about the same thing sit close together.
- Index. Store the embeddings, usually alongside a keyword index, in a vector store the search can query quickly.
- Retrieve. When a question arrives, embed it the same way and fetch the chunks closest in meaning, often combined with keyword matches.
- Augment and generate. Put the top few chunks into the prompt with the question and an instruction to answer only from them.
The first four steps happen ahead of time and again whenever the documents change. The last two happen on every question, in a fraction of a second, before the model writes a word.
Chunking and embeddings, without the math
Chunk size is a tradeoff. Too small, and a passage loses the sentence that explained it; too large, and the search returns pages where one line mattered. Chunks usually overlap so an idea split across a boundary survives in one of them. OpenAI’s retrieval guide (opens in a new tab) shows typical defaults: files added to its vector stores are automatically chunked, embedded and indexed, in 800-token chunks with a 400-token overlap unless you choose otherwise.
Embeddings are what let retrieval find meaning rather than matching words. OpenAI’s guide gives the example of the question “When did we go to the moon?”: the most relevant passage, about the first lunar landing in July 1969, shares none of the question’s words, and semantic search still ranks it first. Keyword search still matters for exact terms such as a policy number, a product code or a person’s name, which is why most systems combine the two. You do not need an embedding model from the same company as your chat model: Anthropic, for example, does not offer its own and points developers to Voyage AI.
A worked example: the handbook assistant
A 60-person company in Columbus, Ohio, puts its employee handbook, benefits guide and holiday calendar behind an internal assistant. An employee asks: “Can I carry unused PTO into next year?” Retrieval returns two chunks, one from the handbook’s time-off section and one from a policy update. What the model actually receives looks like this:
Answer the employee's question using only the sources below. Cite the source id after each sentence, like [S1]. If the sources do not answer it, say "The handbook doesn't say" and suggest contacting HR. Do not guess. <source id="S1" title="Employee Handbook 2026, Time off, p. 14"> Full-time employees may carry over up to 40 hours of unused PTO into the next calendar year. Carried-over hours expire on March 31. </source> <source id="S2" title="Policy update, July 1, 2026"> Part-time employees are not eligible for PTO carryover. </source> Question: Can I carry unused PTO into next year?
A good answer says yes for full-time employees, up to 40 hours, used by March 31 [S1], not for part-time employees [S2], and nothing else. Notice what made that possible: the second chunk was retrieved, the instruction forbade guessing, and every sentence carries a source the employee can open. Take away any one of the three and the answer gets worse in a way the employee cannot see.
Grounding and citations
Grounding means tying an answer to verifiable sources, and citations are how a reader checks it. Asking the model to “cite your sources” in the prompt works up to a point; some APIs now do it structurally. Anthropic’s citations feature (opens in a new tab) returns the exact passages that support each claim, and its documentation says that because the API extracts the cited text directly, citations are guaranteed to point to the documents you provided. Google describes grounding as connecting model output to verifiable sources, with links to them for auditability.
Grounding reduces made-up answers; it does not end them. A model can still misread a passage, merge two that disagree, or answer confidently when retrieval fetched the wrong chunk. NIST’s Generative AI Profile (opens in a new tab) calls confidently stated but false content confabulation, and RAG is a way to make it rarer and easier to catch, not a guarantee against it.
When to use RAG, and when not
- Use RAG when the answers are written down in a large or changing set of documents: policies, product docs, contracts, past tickets. It is also the right choice when answers must cite a source.
- Use a long context window instead when the whole set is small enough to send in full. Google’s guide to long context (opens in a new tab) says many Gemini models take a million tokens or more, but also that accuracy drops when a question depends on several separate facts scattered through that much text, and that you pay for every token on every request unless you cache them.
- Use MCP tools when the data is live or the assistant must act: a task board, a CRM or a ticket queue, where the model should look things up while it works and sometimes change them.
- Use fine-tuning rarely, for behavior rather than facts: a format, a tone or one narrow task at high volume. It bakes a snapshot into the model and cannot cite where an answer came from.
Fine-tuning is also becoming harder to get from the big vendors. As of September 30, 2026, Anthropic’s documentation says the Claude API does not offer fine-tuning, and OpenAI’s model optimization guide (opens in a new tab) says it is winding down its fine-tuning platform, which is closed to new users.
- Choosing by symptom, with a decision order: prompt engineering vs RAG vs fine-tuning.
- Retrieval before the answer versus tools during the work: MCP vs RAG.
- Documents people wrote versus what an agent learned while working: agent memory vs RAG.
A checklist before you trust a RAG system
- Ask ten real questions whose answers you know, including two the documents do not answer. It should say so for those two.
- Check freshness. Change a document and see how long the answer takes to change.
- Check permissions. Retrieval that ignores who is asking will quote documents that person could never open.
- Look at what was retrieved, not only the answer. Most bad answers start with a bad search.
- Require citations, and open a few every week.
- Treat retrieved text as data, not instructions. A document can contain text written to steer the model, as indirect prompt injection explains.
- Keep one owner for the document set. RAG is only as good as what it retrieves, and stale documents produce confident stale answers.
Where fenbs fits
fenbs is a task board, not a RAG system: it has no vector index and does not answer questions from your documents. What it offers an assistant is the other half. Over MCP, fenbs_get_context hands a connected assistant the board’s rules from the Decisions and rules page and its AI context notes at the start of a session, and fenbs_search looks up tasks live, by their words and refs, so results are never out of date. Building and testing a RAG system is work in its own right, and it fits a board well: file each retrieval failure as a bug with the question in the note, put the fix you will try in the plan, and record the test status and test notes when you rerun the questions.
Related
Why a model forgets between sessions: AI context windows explained. Writing the instruction part of a RAG prompt: prompt engineering. Using ChatGPT with company documents safely: ChatGPT for business. What MCP is: the glossary.