Published 27 September 2026 · Definitions, context-window figures and benchmark methods checked against the linked sources on 26 September 2026
RAG, short for retrieval-augmented generation, is a technique that lets a large language model answer from information it was never trained on: when a question arrives, the application searches a knowledge source for the most relevant passages, places them in the prompt next to the question, and the model writes its answer from that material. The model's weights never change. What changes is what the model can see on that one call, which is why RAG became the standard way to put private documents, current policies or a product manual in front of a general-purpose model.
Generation is the easy half of that sentence. Retrieval is where the engineering lives, and it is less settled than most explanations admit. When we ran 50,000 documents and 5,000 queries through keyword search and vector search, vector search won the five-dataset average by 0.6320 to 0.5325 on MRR@10, then lost the average of the other four once the single natural-language-to-code dataset was taken out. A RAG system is only as good as the passage it hands the model, and how you find that passage is a measurable decision rather than a default.
What RAG stands for, and where the name came from
RAG stands for Retrieval-Augmented Generation: retrieve relevant text, augment the prompt with it, generate the answer. The term comes from a 2020 paper by Patrick Lewis and colleagues from Facebook AI Research, University College London and New York University, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, which paired what the authors called parametric memory (a pre-trained sequence-to-sequence model) with non-parametric memory (a dense vector index of Wikipedia). Lewis later told NVIDIA that the team would have put more thought into the name had they known how widely it would spread.
Keep one detail from that paper in mind for later. The original design already treated the retrieval index as memory, but a read-only one, and that distinction decides exactly where RAG stops being enough.
Why RAG exists
A language model knows what was in its training data up to a cutoff date, and nothing after it. It has never seen your refund policy or last week's release notes, and when asked about them it tends to produce a fluent answer anyway, which is the failure people call hallucination. Retraining or fine-tuning a model every time a document changes is slow and expensive, and it still would not let the model point to where an answer came from.
Retrieval takes care of both problems at request time. Fresh or private text arrives in the prompt on every call, and because the application knows which passages it supplied, the answer can cite them so a person can check. AWS and IBM describe the same trade in their explainers: extend a general model to your domain without retraining it.
What RAG does not do is guarantee correctness. It reduces hallucination by giving the model better material; the model can still misread that material, and it can answer confidently when retrieval returned nothing useful. Wikipedia's entry collects documented cases, and IBM's own explainer says plainly that RAG "cannot make a model error-proof."
How RAG works, step by step
A RAG system runs in two phases at different times. Indexing happens ahead of time, whenever the source documents change; retrieval and generation happen on every request.
Indexing, ahead of time
- Collect the sources. PDFs, wiki pages, support tickets, database rows, anything the model should be able to answer from.
- Split them into chunks. A chunk is the unit retrieval returns, often a few hundred tokens. Chunk size is a real tuning knob: a large chunk carries its surrounding context but matches queries loosely, while a small chunk matches precisely and can cut a fact away from the sentence that qualifies it.
- Index each chunk. For vector search, an embedding model turns the chunk into a list of numbers that sits close to other text with similar meaning, and the vectors go into a vector database or into PostgreSQL with the pgvector extension. For keyword search, the chunk's terms go into an inverted index, the same structure a search engine uses.
Retrieval and generation, per request
- Retrieve. The question becomes a query of the same kind (an embedding, a set of terms, or both), and the index returns the top candidate chunks.
- Rerank, if you need it. A slower, more precise model re-scores those candidates against the question and keeps the best few.
- Augment the prompt. The application assembles instructions, the retrieved passages and the question into one prompt.
- Generate. The model answers from that prompt, ideally citing the passages it used.
Step 6, the augmented prompt, is plainer than its name suggests. A version close to the template Pinecone publishes looks like this:
Answer the QUESTION using only the CONTEXT below.
If the CONTEXT does not contain the answer, say you do not know.
CONTEXT:
[1] <retrieved passage>
[2] <retrieved passage>
QUESTION: <the user's question>
That second line earns its place, because it gives the model permission to report that retrieval came back empty instead of filling the gap with something plausible.
Put together, the online path is a handful of lines. This sketch uses no particular library; index, rerank and llm stand in for whatever you run.
def answer(question: str, k: int = 5) -> str:
candidates = index.search(question, top_k=20) # keyword, vector, or hybrid
passages = rerank(question, candidates)[:k] # optional second pass
context = "\n".join(f"[{i + 1}] {p.text}" for i, p in enumerate(passages))
prompt = (
"Answer the QUESTION using only the CONTEXT. "
"If the CONTEXT does not contain the answer, say you do not know.\n\n"
f"CONTEXT:\n{context}\n\nQUESTION: {question}"
)
return llm.generate(prompt)
Each piece of that pipeline is a separate choice:
| Component | What it does | Common choices |
|---|---|---|
| Knowledge source | The material RAG answers from | Files, wikis, tickets, database rows |
| Chunker | Splits documents into retrievable units | Fixed size, by heading, by paragraph |
| Indexer | Makes chunks searchable | An embedding model for vectors, an inverted index for keywords |
| Store | Holds the index | A vector database, PostgreSQL with pgvector, a search library such as Tantivy |
| Retriever | Finds candidate chunks for a query | Vector similarity, keyword scoring such as BM25, or both |
| Reranker | Re-scores the candidates | A cross-encoder or a hosted reranking model |
| Orchestrator | Builds the prompt and calls the model | Your own code, LangChain, LlamaIndex |
| Generator | Writes the answer | Any capable LLM |
Does RAG need a vector database?
No. Most RAG explainers go straight from "retrieve" to "embed and query a vector database", as if the two were the same thing. Retrieval can be keyword search, vector search or a hybrid of both, and which one wins depends on how your users phrase questions compared with how your documents are written.
We measured it rather than assuming it. The study ingested 50,000 documents, 10,000 from each of five public datasets, and ran 5,000 queries against two retrievers: Tantivy 0.22.0 on its default analyzer for exact-match keyword search, and ChromaDB with the all-MiniLM-L6-v2 embedding model at 384 dimensions for vector search, with no chunking on either side. It ran on 14 January 2026 on an Apple M4 with 16GB of RAM as a single run, scored on MRR@10 by exact document-ID match against each dataset's ground truth, with no LLM judge anywhere in the pipeline.
| Dataset | What it tests | Keyword (MRR@10) | Vector (MRR@10) | Result |
|---|---|---|---|---|
| CodeXGLUE | Natural language to code | 0.2901 | 0.9143 | Vector |
| MS MARCO | Real Bing queries | 0.4035 | 0.5225 | Vector |
| SQuAD | Wikipedia question answering | 0.6048 | 0.6136 | Tie, inside the noise of a single run |
| HotpotQA | Multi-hop reasoning | 0.5494 | 0.4953 | Keyword |
| SciQ | Science exam questions | 0.8145 | 0.6142 | Keyword |
| Mean, all five | 0.5325 | 0.6320 | Vector | |
| Mean, without CodeXGLUE | 0.5931 | 0.5614 | Keyword |
Read the last two rows together. Vector search is not generally better at retrieval; it is decisively better at bridging a vocabulary gap, where a query like "sort a list" has to find a function called bubble_sort, and that one case carries the average. Where the words in the question are the words in the document, as with scientific terms, product names or error codes, exact matching locks on and vector search drifts toward text that is merely nearby in meaning.
Cost points the same way. Indexing the 50,000 documents took 2.11 seconds with keyword search and 161.6 seconds with embedding generation on the same machine, a factor of 76.4. For an agent that has to read a repository and act on it within a single turn, that is the difference between instant and a noticeable wait; for a corpus indexed once and queried for months, it amortises away.
One caveat the original write-up states itself: all-MiniLM-L6-v2 is a small, older embedding model, and a stronger one would very likely lift the vector scores, so the table describes a common default configuration rather than the best vector setup available. The MTEB leaderboard is the place to compare current embedding models on a task shaped like yours.
Many systems settle the question by running both retrievers and merging the results, which is called hybrid search, and then reranking the merged list. Anthropic reported that combining contextual embeddings with contextual BM25 keyword matching cut its top-20 retrieval failure rate by 49%, and by 67% once a reranker was added, on its own test sets.
RAG, fine-tuning, or a longer prompt
Four approaches can give a model knowledge or behaviour it did not ship with, RAG among them, and each one changes something different. Databricks lays the options out as a ladder, summarised here:
| Approach | What changes | What it needs | Best for | Main cost |
|---|---|---|---|---|
| Prompt engineering | The instructions | Nothing extra | Steering format and tone quickly | Least control |
| RAG | What the model sees on each request | A knowledge source and an index | Facts that change, or that are private | Extra prompt tokens and a retrieval step per call |
| Fine-tuning | The model's weights | Thousands of domain or instruction examples | Behaviour, output format, domain vocabulary | Training compute, repeated when the data shifts |
| Pretraining | Everything | Billions to trillions of tokens | A model for a genuinely new domain | Very high |
One rule survives most real decisions: facts that change belong in retrieval, and behaviour belongs in fine-tuning. A model fine-tuned on your March price list will quote March prices with full confidence in September, while a RAG system reads the current list on every call. The two combine well, since a model fine-tuned to follow your output format can still read retrieved passages.
A newer objection is the context window itself. Anthropic's documentation lists Claude Opus 5.5 and Claude Sonnet 5 with 1M-token context windows, enough to paste a large manual into every prompt. Two things keep RAG in the picture. Every token you send is billed on every call, so a million-token prompt is a million-token cost per turn. And the same Anthropic page states that as token count grows, accuracy and recall degrade, a phenomenon it calls context rot; the Lost in the Middle study found models use information at the start or end of a long input better than information buried in the middle. A long window is a good place to reason over the right material, and RAG is how the right material gets there. Our post on why a bigger context window does not settle this covers the economics and the latency, and our token cost breakdown for long conversations shows where caching stops helping.
Good fits are wherever the answer lives in a document someone maintains: a support assistant answering from product docs and policies, internal search over HR and engineering wikis, a research assistant over a paper collection, a sales assistant quoting the current price sheet. If the answer would be the same for every user who asks, and it is written down somewhere, RAG is the right starting point.
Semantic search and RAG are often used as synonyms, and they are not. Semantic search is a retrieval technique that finds passages by meaning; RAG is the pattern of handing whatever was retrieved to a generator. You can run semantic search without RAG (a search box that returns links) and RAG without semantic search (keyword retrieval feeding a model).
Where RAG breaks
Retrieval misses are the most common failure, and they are silent. If the right passage is not in the top results, the model raises no error; it answers from whatever it was given. Two frequent causes are chunking, when a fact and its qualifier land in different chunks, and a stale index, when a document changed and nobody re-embedded it. AWS recommends updating documents and their embeddings asynchronously, on a schedule or as changes arrive, which is a reminder that freshness is a pipeline you operate rather than a property you get for free.
Multi-hop questions break similarity retrieval more fundamentally. In the HotpotQA portion of our study, questions such as which of two magazines was founded first need two documents, and vector search reliably found the document for the first entity while missing the bridge document for the second, because that second document did not resemble a query that was mostly about the first. No reranker fixes this, since both retrievers score each document independently against the query. Answering it requires storing the relationship between documents, which is why graph approaches exist; Microsoft's GraphRAG paper builds an entity graph and community summaries so that questions spanning a whole corpus can be answered, reporting gains over a conventional RAG baseline on those global questions. Our GraphRAG glossary entry covers the pattern.
Access control has to live inside the retrieval query. A common pattern is one shared index with a tenant or user identifier in the query filter, and when that filter is dropped or malformed, an ordinary application throws an error while a RAG system returns a fluent answer assembled from someone else's documents, with nothing in monitoring to show it happened. IBM also points out that a breached vector store can expose the original text, so the index deserves the same protection as the source data.
Evaluate retrieval and generation separately. For retrieval, collect real questions with known source documents and measure whether the right document ranks near the top, with recall@k or MRR. For generation, check whether answers are grounded in the retrieved passages and whether they are correct, which needs a human reviewer or a carefully chosen model judge. Keeping the two apart tells you which half to fix: a correct passage with a wrong answer points at the prompt or the model, and a wrong passage points at retrieval. Our own study carries the same caveat, since MRR@10 measures where the right document ranks and not whether the final answer was right.
Agentic RAG is the newer shape of the same idea. Instead of one fixed retrieve-then-generate pass, an agent decides whether to retrieve at all, rewrites the query, can call more than one retrieval tool, and checks what came back before it answers, as Pinecone's explainer describes. It recovers from some misses by trying again. It does not change what RAG fundamentally is, a read of a corpus somebody else wrote.
Where RAG stops and memory starts
Every RAG system described so far reads from a corpus that somebody else wrote and that does not change during the conversation. Nothing about the user, the conversation or what the agent did is written back. For a documentation assistant that is exactly right, since "what is the refund policy" should get the same answer for everyone.
For an agent that talks to the same person again and again, it is the missing half. A customer who said in March that they live in Chicago and in June that they moved to Berlin needs a system that decides what is worth keeping from those conversations, recognises that the new fact replaces the old one, keeps the change traceable, and keeps that person's facts out of every other user's session. A vector index over past transcripts returns both lines as equally relevant text with no view on which is current. That is the line between the two: retrieval is a read path, and memory is a read path plus a write path, with rules for what to keep, what to update, what to retire and who may see it.
Maximem Synap is built on the second model. Rather than storing transcripts and searching them later, it reads each conversation as it arrives, extracts self-contained statements such as "The user lives in Berlin" with a type, a confidence score, a timestamp and a scope, and when a new statement replaces an old one it marks the old one historical and links it to its replacement instead of filing both side by side. Retrieval is still a step inside Synap; what changes is the thing being retrieved. How to build this yourself, and when RAG alone is enough, is covered in how to give an LLM long-term memory, and the wider argument is in what agentic context management is and why AI forgets.
RAG made it cheap to give a model your documents. Giving an agent a reliable record of the people it serves is the next step, and it is the one that stops a customer from repeating themselves on the third call.
Frequently asked questions
What does RAG stand for in AI?
RAG stands for Retrieval-Augmented Generation, a technique in which an application retrieves relevant passages from a knowledge source and adds them to a language model's prompt before the model generates its answer.
Does RAG require a vector database?
No. RAG needs a retriever, and the retriever can use keyword search, vector search or both. In Maximem's study of 50,000 documents and 5,000 queries, vector search won the five-dataset average on MRR@10, but keyword search won the average of the remaining four datasets once the natural-language-to-code dataset was removed, 0.5931 to 0.5614, and indexed 76.4 times faster.
Does RAG stop hallucinations?
RAG reduces hallucinations by giving the model relevant source text, and it lets answers cite that text. It does not eliminate them: a model can still misread retrieved passages, and it can answer confidently when retrieval returned nothing useful.
Do I still need RAG with a 1M-token context window?
Usually, yes. Anthropic lists 1M-token context windows for Claude Opus 5.5 and Claude Sonnet 5, and the same documentation notes that accuracy and recall degrade as token count grows. Sending a whole corpus on every call also bills every token on every call. RAG sends the relevant passages, and the long window gives the model room to reason over them.
Is RAG the same as agent memory?
No. RAG reads from a corpus written elsewhere and writes nothing back, so it answers the same way for every user. Agent memory adds a write path that decides what to keep from each conversation, updates facts when they change, and scopes them to the right person. Memory systems use retrieval internally, but retrieval on its own is not memory.
Who invented RAG?
Patrick Lewis and co-authors from Facebook AI Research, University College London and New York University coined the term in a May 2020 paper, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", which paired a pre-trained sequence-to-sequence model with a dense vector index of Wikipedia.
What is agentic RAG?
Agentic RAG puts an agent in charge of retrieval: it decides whether to search, rewrites the query, can call several retrieval tools, and checks the results before answering. It recovers from some retrieval misses by trying again, but it still reads a fixed corpus rather than remembering anything about the user.
Sources, retrieved 26 September 2026: Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv 2005.11401); NVIDIA, What Is Retrieval-Augmented Generation; AWS, What is RAG; IBM, What is RAG; Databricks, What is Retrieval Augmented Generation; Pinecone, Retrieval-Augmented Generation; Google Cloud, What is RAG; Wikipedia, Retrieval-augmented generation; Anthropic, Context windows; Anthropic, Contextual Retrieval; Liu et al., Lost in the Middle (arXiv 2307.03172); Edge et al., From Local to Global: A Graph RAG Approach (arXiv 2404.16130); MTEB leaderboard; pgvector; Maximem, File Search vs Vector Search for RAG: 50,000 Documents, 5,000 Queries; Maximem, Build vs buy agent memory.



