What Is RAG? Retrieval-Augmented Generation Explained Simply
RAG lets an AI answer from your own documents by finding the relevant passages first. How it works, embeddings, chunking, costs and what to ask your AI.
Mythex Team · · 6 min read
RAG (retrieval-augmented generation) is a way to make an AI model answer questions using your own information. When someone asks a question, your app first searches your documents for the most relevant passages, then hands those passages to the model along with the question. The model writes its answer from that text instead of relying only on what it happened to learn during training.
Why RAG matters when you build with AI
A large language model knows a lot about the world in general, but nothing about your business: your prices, your policies, your product manual, last week's support tickets. It also has a training cutoff, so it doesn't know about anything recent. Ask it anyway and it may produce a confident, plausible and wrong answer.
You have three broad options to fix that:
| Approach | What it does | Good for |
|---|---|---|
| Paste everything into the prompt | Send all your documents with every question | Small amounts of text — a few pages |
| RAG | Search your documents, send only the relevant parts | Larger or changing collections: help centres, manuals, notes |
| Fine-tuning | Train the model further on examples | Tone, format and repeated tasks — not facts that change |
Modern models accept long prompts, so for a short FAQ the first option is often fine. RAG earns its place when you have more text than you want to send every time — because every word you send costs money and time (see how LLM APIs charge).
If you're building a "chat with our docs" feature, a support assistant or an internal knowledge tool, you're almost certainly building some form of RAG.
An everyday analogy
Imagine an open-book exam. A student who has to answer from memory may misremember or bluff. A student allowed to flip to the right pages first, read them and then answer is far more reliable — as long as they find the right pages.
- The student is the language model.
- The textbook is your documents.
- Flipping to the right pages is retrieval.
- Writing the answer from those pages is generation.
RAG is the open-book version. And, as in the exam, the quality of the answer depends heavily on whether the right pages were found.
How RAG works, step by step
A typical RAG setup has two phases.
Preparing your documents (done ahead of time):
- Collect the source material — help articles, PDFs, policies, product data.
- Chunk it: split long documents into smaller pieces, often a few paragraphs each, so the search can return just the relevant part.
- Embed each chunk: send it to an embedding model, which turns the text into an embedding — a long list of numbers that captures its meaning. Texts about similar things end up with similar numbers.
- Store the chunks and their embeddings in a database that can search by similarity (often called a vector database or vector index).
Answering a question (done every time):
- Embed the question with the same embedding model.
- Retrieve the chunks whose embeddings are closest to the question's — say, the top five.
- Build the prompt: instructions, the retrieved chunks, and the question.
- Generate: the language model writes an answer based on those chunks, ideally citing which ones it used.
Searching by meaning is what makes this work: a question about "getting my money back" can find a chunk titled "Refund policy" even though they share no words.
A worked example: a gym's help assistant
Say a gym has 40 help articles and wants a chat assistant on its website. A member asks: "Can I pause my membership while I'm travelling?"
The app embeds the question and retrieves the three closest chunks. One of them reads:
Membership freeze: Members on monthly plans can freeze their
membership for up to 2 months per year. Request it at the front
desk or in the app at least 7 days before the freeze starts.
The prompt sent to the model then looks something like this:
You are the help assistant for Northside Gym.
Answer only from the sources below. If the answer isn't in them,
say you don't know and suggest contacting the front desk.
Mention which source you used.
[Source 1: Membership freeze] Members on monthly plans can freeze...
[Source 2: Cancelling your membership] ...
[Source 3: Guest passes] ...
Question: Can I pause my membership while I'm travelling?
And the model answers: "Yes — on a monthly plan you can freeze your membership for up to two months a year. Request it at the front desk or in the app at least 7 days before (source: Membership freeze)."
The gym never retrained anything. When it changes the freeze policy, it updates the article, the app re-embeds that chunk, and the next answer reflects it.
Key terms
| Term | Meaning |
|---|---|
| LLM | Large language model — the AI that writes the answer. See what an LLM is. |
| Embedding | A list of numbers representing the meaning of a piece of text. |
| Embedding model | A model that produces embeddings. Usually separate from, and much cheaper than, the chat model. |
| Chunk | A piece of a document, sized so it can be retrieved on its own. |
| Vector database / vector search | Storage that finds items with similar embeddings. Some regular databases offer it too — Postgres, for example, through the pgvector extension. |
| Top-k | How many chunks are retrieved per question. |
| Hybrid search | Combining meaning-based search with ordinary keyword search, which helps with exact terms like product codes. |
| Context window | The maximum amount of text a model can take in one request. |
| Grounding / citations | Tying the answer to specific retrieved sources and showing them. |
Common misconceptions
- "RAG trains the model on my data." It doesn't. Nothing about the model changes; it just reads extra text at question time.
- "RAG means no more hallucinations." It means fewer. If retrieval misses the right chunk, or the answer isn't in your documents, the model may still guess. Instruct it to say "I don't know" and test with questions whose answers you know.
- "I need a special vector database." Often you don't. For modest amounts of data, a vector extension on the database you already have, or even plain keyword search, can be enough.
- "Bigger chunks are better." Too big and you send lots of irrelevant text (more cost, more confusion); too small and a chunk loses the context that made it meaningful. Expect to adjust.
- "Just upload everything." Outdated, duplicated or contradictory documents produce contradictory answers. Cleaning the source material is often the biggest improvement you can make.
- "Anyone can see any document." Only if you build it that way. If some documents are private to certain users, the search itself must filter by who is asking — otherwise the assistant can quote one customer's data to another.
What it costs
A RAG feature usually has three running costs: embedding your documents (small, and mostly one-off), embedding each question (tiny), and the language model call (the main cost, which grows with how many chunks you send). Plus storage in your database. Your AI provider bills these to your own account, so set a spending limit there.
What to ask your AI builder
- "Before building RAG, check: is our content small enough to just include in the prompt?"
- "Split the help articles into chunks of a few paragraphs, keep each chunk's title and source link, and store embeddings in our database."
- "Use hybrid search — meaning-based plus keyword — so product codes still match."
- "Tell the model to answer only from the retrieved sources, say when it doesn't know, and show the sources under each answer."
- "Only search documents the signed-in user is allowed to see."
- "Re-embed a document automatically whenever it's edited."
- "Keep the AI provider key on the server in a secret, never in the browser."
- "Log each question, the chunks retrieved and the answer so I can review bad answers."
RAG in Mythex
Mythex doesn't include a ready-made knowledge-base or RAG product. What it does is write the code for one inside your app: you bring your documents and your own AI provider key, stored as a project secret, and the agent builds the retrieval and chat parts, using the project's Postgres database for storage. Provider usage bills your provider account, not Mythex credits. The docs suggest starting with one simple AI feature and adding RAG once that works — see building an AI app, and our guides on building an AI chatbot app and adding AI features.
Questions
What is RAG in simple terms?
RAG (retrieval-augmented generation) is a way to make an AI model answer from your own information. Before the model replies, your app searches your documents for the passages most relevant to the question and includes them in the prompt, so the answer is based on that text rather than only on what the model learned in training.
Does RAG train the AI on my data?
No. RAG doesn't change the model at all. It looks up relevant text at the moment a question is asked and passes it to the model along with the question. Update a document and the next answer can use the new version straight away.
Is RAG the same as fine-tuning?
No. Fine-tuning changes a model's behaviour by training it further on examples, which is slower and better for style or format. RAG gives the model facts to read at question time, which is better for knowledge that changes or must be quoted accurately.
Does RAG stop AI hallucinations?
It reduces them but doesn't eliminate them. If the search finds the wrong passages, or the answer isn't in your documents, the model can still guess. Good RAG apps tell the model to say when it doesn't know and show the sources they used.