What is RAG?
A short learning video. No voice, only titles and screens, with music.
What this video shows
The answer comes first, in two lines. The AI looks in your files first. Then it answers from the page it finds. On screen, a gold line goes from the chat to a shelf of your files, and a page comes back.
Then one example. Someone asks: “Can a customer return a chair late?” A gold glint sweeps along the shelf and stops on one file, Returns. A page from it slides into the chat, one line in the page glows, and the answer appears at the same moment: “Yes. Returns are free for a month.” You can see exactly where the answer came from.
The last part is what to do. Give it your own files. Add a file when things change: a new one, Sofas, is dropped on the shelf, and from then on it is part of what the AI can look up. And ask it to show the page it used: it answers “Here it is: the Returns file, line three.” and brings the page back.
That is RAG in plain words. The letters stand for retrieval-augmented generation, which is the same thing in longer words. It is how almost every AI app I build for daily work answers questions about a business’s own prices, policies, orders and documents. The model stays as it is; the shelf is yours.
What you will learn
- Why the AI does not know your business: it was trained on the world, and your files were never in it.
- The three moves of RAG: it looks in your files, it puts the right page into the chat, and it answers from that page.
- What to do: give it your files, add new ones when things change, and ask it to show the page it used.
- Why adding a file is enough: nothing is retrained; the next lookup simply finds it.
The numbers behind the picture
Nothing with a digit is on screen. The dates and numbers belong here.
- 2025: Half of production AI designs used RAG in 2024 (51%, up from 31% the year before), and a year later it was still the second most used method after prompt design. With the source in front of it, the best models made something up in under one summary in a hundred (October 2025). The weak point is the lookup: one study measured a missed or wrong page as the single largest cause of RAG errors, 32.5%. Anthropic’s “contextual retrieval” method cut missed lookups by 49%, and by 67% with a second ranking pass. (Menlo Ventures, 2024 and Dec 2025; Vectara leaderboard, 16 Oct 2025; FAIR-RAG, Oct 2025; Anthropic, Sep 2024)
- 2026: Google’s Gemini Notebook, a “chat with your documents” tool, reached 30 million people and 600,000 organisations in July. Google’s and OpenAI’s developer products now offer the whole pipeline as one feature (Gemini File Search, Sep 2026; Vertex RAG Engine, Oct 2026; OpenAI file search with 800-token chunks). Vectara’s 2026 leaderboard uses a harder judge; the best model is at 1.8%, so the “under one in a hundred” line belongs to the 2025 test. (Google, 16 Jul 2026; Google and OpenAI docs, 2026; Vectara, Sep 2026)
- 2027, my guess: A shelf beside the chat becomes the normal way a business app answers. The question stops being “can it look things up?” and becomes “did it find the right page, and does it say so when it did not?” The apps that show their page will be the ones people trust.
Questions people ask
What does RAG stand for?
Retrieval-augmented generation. In plain words: the AI retrieves the right pages from your files, adds them to the chat, and generates the answer from them. The idea comes from a 2020 research paper by Lewis and others; today every big AI vendor offers it as a built-in feature.
Why not just paste all my documents into the chat?
For a small set, you can; Anthropic’s own advice is that below a certain size you simply put it all in. Past that, the AI’s working window fills up and its accuracy falls, which studies in 2025 and 2026 measured across many models. RAG keeps the chat small by bringing in only the pages that match the question.
Does RAG stop the AI from making things up?
It cuts it hard, because the answer is tied to a page you can open. With the source given, the best models made something up in under one summary in a hundred on the October 2025 test. It does not fix everything: if the lookup brings the wrong page, the answer is wrong with confidence. That is why a good RAG app shows its page, and says “not in your files” when nothing fits. The video “Why does AI sometimes make things up?” shows the other half.
What does a small business need to start?
The files, in one place, and a clear rule for what the AI may answer from them. Prices, policies, product pages, order records and how-to notes are the usual shelf. The app does the chunking and matching; the owner’s job is to keep the shelf current, which is the “add a file” scene in the film.
Further reading
- Lewis and others, “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, NeurIPS 2020. Where the name comes from. arxiv.org
- Anthropic, “Introducing Contextual Retrieval”, September 2024. The 49% and 67% figures, and “just put it all in the prompt” for small sets. www.anthropic.com
- Anthropic developer docs, “Citations”. The exact passages behind each claim. platform.claude.com
- Google, Gemini API, “Grounding” and “File Search” docs (2026). What grounding means, and the built-in pipeline. ai.google.dev
- Google Cloud, Vertex AI RAG Engine overview (updated October 2026). cloud.google.com
- OpenAI developer docs, “File search”. Chunking defaults and citations. developers.openai.com
- Anthropic, “Effective context engineering for AI agents”, September 2025. Why you cannot paste everything: context rot and just-in-time retrieval. www.anthropic.com
- Vectara, Hallucination Leaderboard, October 2025 and September 2026 tables. github.com
- Menlo Ventures, “The State of Generative AI in the Enterprise”, 2024 and December 2025. The RAG adoption figures. menlovc.com
Every AI app I build for a business answers the way this film shows: from the business’s own shelf, with the page in view, and “not in your files” when nothing fits (how I can help). If your team wants answers it can check, tell me about the work.
Transcript
Every line of text shown on screen, in order. Nothing is spoken. Each label stays at the bottom of the screen for at least four seconds.
Opening
- Logo: vai SoftLab (small, no title line)
Scene one: the question
Alone on a dark screen, the biggest text in the film:
- What is RAG?
Scene two: the answer
A chat window on the left with a small robot face beside it. On the right, a shelf of small file icons under a gold tag. Three files have titles.
- Chat
- Your files
- Returns
- Delivery
- Prices
A gold line goes from the robot to the shelf; one file glows.
Label: The AI looks in your files first.
A page slides from that file into the chat, and a short gold answer appears under it.
Label: Then it answers from the page it finds.
Scene three: one example
Typed, letter by letter, then sent:
- Can a customer return a chair late?
A gold glint sweeps along the shelf and stops on Returns; the file rises and turns gold.
Label: It finds the Returns file.
A page slides from the file into the chat. One line in the page glows.
Label: It brings the right page in.
A gold answer appears, and the glowing line pulses at the same moment:
- Yes. Returns are free for a month.
Label: The answer comes from that page.
Scene four: what to do
A hand drags two more files onto the shelf.
Label: Give it your own files.
The hand drops a new file into the top row:
- Sofas
Label: Add a file when things change.
A typed line:
- Show me the page you used.
A gold answer:
- Here it is: the Returns file, line three.
The Returns page slides back into the chat with its glowing line.
Label: Ask it to show the page.
Scene five: to remember
Dark screen. Two lines:
- RAG: it looks in your files,
- then answers from the page.
Closing
- Logo: vai SoftLab
- vaisoftlab.com
Have an app idea, a stuck app, or a daily job you want an AI agent to do?
Write three lines. I reply within one working day with a plain answer.


