The simple explanation
RAG — retrieval-augmented generation — is an open-book setup with two jobs. First, retrieve passages relevant to the question from a knowledge store: manuals, articles, notes. Then give those passages to the language model along with the question, so it answers with the evidence in front of it. Retrieval can use keyword matching, embeddings, or both.
The key point: RAG does not train the model on your documents. The model never memorizes your handbook; it reads the relevant pages at answer time, like a student allowed notes in an exam. Remove the documents and the answers go back to pure guessing.
In practice there’s a step before any of this: your documents get chunked — split into short passages — and stored in a searchable index. Good chunking keeps each passage self-contained; bad chunking splits an answer across two chunks and the retriever finds neither. Most of RAG’s real-world reliability lives in these unglamorous details: how the text was split, how the search ranks, and how many passages get handed to the model.
See this idea move.
The RAG Visualizer experiment walks through this concept step by step — press run and watch it happen. Everything is simulated in your browser; no real AI runs.
A concrete example
Ask the RAG Visualizer about returns in the tiny shop handbook. The system finds the matching passage and the answer quotes it — answers with receipts. Now remove the source and ask again: the honest system admits it lacks evidence instead of inventing a policy. That’s the whole promise: grounded answers when the evidence exists, and an honest “I don’t know” when it doesn’t.
Where people get misled
The big one: “RAG eliminates hallucinations.” It reduces them; it doesn’t eliminate them. Retrieval can pull the wrong passage, a stale passage, or a passage that almost answers. And even with the right passage, the model can misquote it, overclaim beyond it, or blend it with its own guesses. Garbage in is still garbage out — if your documents are wrong or outdated, RAG serves wrong answers with citations.
The honest limits
Check the sources, the dates, and the permissions. RAG systems can retrieve documents the asker was never meant to see, so access control matters. And generated answers can still misrepresent good passages — important claims get checked against the source itself, not against the model’s summary of it.