Guides / Concepts explained

What Is RAG? Retrieval-Augmented Generation Explained Simply

RAG lets an AI answer from your own documents by finding the relevant passages first. How it works, embeddings, chunking, costs and what to ask your AI.

Mythex Team · 2026-09-29 · 6 min read

RAG (retrieval-augmented generation) is a way to make an AI model answer questions using your own information. When someone asks a question, your app first searches your documents for the most relevant passages, then hands those passages to the model along with the question. The model writes its answer from that text instead of relying only on what it happened to learn during training.

Why RAG matters when you build with AI

A large language model knows a lot about the world in general, but nothing about your business: your prices, your policies, your product manual, last week's support tickets. It also has a training cutoff, so it doesn't know about anything recent. Ask it anyway and it may produce a confident, plausible and wrong answer.

You have three broad options to fix that:

ApproachWhat it doesGood for
Paste everything into the promptSend all your documents with every questionSmall amounts of text — a few pages
RAGSearch your documents, send only the relevant partsLarger or changing collections: help centres, manuals, notes
Fine-tuningTrain the model further on examplesTone, format and repeated tasks — not facts that change

Modern models accept long prompts, so for a short FAQ the first option is often fine. RAG earns its place when you have more text than you want to send every time — because every word you send costs money and time (see how LLM APIs charge).

If you're building a "chat with our docs" feature, a support assistant or an internal knowledge tool, you're almost certainly building some form of RAG.

An everyday analogy

Imagine an open-book exam. A student who has to answer from memory may misremember or bluff. A student allowed to flip to the right pages first, read them and then answer is far more reliable — as long as they find the right pages.

  • The student is the language model.
  • The textbook is your documents.
  • Flipping to the right pages is retrieval.
  • Writing the answer from those pages is generation.

RAG is the open-book version. And, as in the exam, the quality of the answer depends heavily on whether the right pages were found.

How RAG works, step by step

A typical RAG setup has two phases.

Preparing your documents (done ahead of time):

  1. Collect the source material — help articles, PDFs, policies, product data.
  2. Chunk it: split long documents into smaller pieces, often a few paragraphs each, so the search can return just the relevant part.
  3. Embed each chunk: send it to an embedding model, which turns the text into an embedding — a long list of numbers that captures its meaning. Texts about similar things end up with similar numbers.
  4. Store the chunks and their embeddings in a database that can search by similarity (often called a vector database or vector index).

Answering a question (done every time):

  1. Embed the question with the same embedding model.
  2. Retrieve the chunks whose embeddings are closest to the question's — say, the top five.
  3. Build the prompt: instructions, the retrieved chunks, and the question.
  4. Generate: the language model writes an answer based on those chunks, ideally citing which ones it used.

Searching by meaning is what makes this work: a question about "getting my money back" can find a chunk titled "Refund policy" even though they share no words.

A worked example: a gym's help assistant

Say a gym has 40 help articles and wants a chat assistant on its website. A member asks: "Can I pause my membership while I'm travelling?"

The app embeds the question and retrieves the three closest chunks. One of them reads:

Membership freeze: Members on monthly plans can freeze their
membership for up to 2 months per year. Request it at the front
desk or in the app at least 7 days before the freeze starts.

The prompt sent to the model then looks something like this:

You are the help assistant for Northside Gym.
Answer only from the sources below. If the answer isn't in them,
say you don't know and suggest contacting the front desk.
Mention which source you used.

[Source 1: Membership freeze] Members on monthly plans can freeze...
[Source 2: Cancelling your membership] ...
[Source 3: Guest passes] ...

Question: Can I pause my membership while I'm travelling?

And the model answers: "Yes — on a monthly plan you can freeze your membership for up to two months a year. Request it at the front desk or in the app at least 7 days before (source: Membership freeze)."

The gym never retrained anything. When it changes the freeze policy, it updates the article, the app re-embeds that chunk, and the next answer reflects it.

Key terms

TermMeaning
LLMLarge language model — the AI that writes the answer. See what an LLM is.
EmbeddingA list of numbers representing the meaning of a piece of text.
Embedding modelA model that produces embeddings. Usually separate from, and much cheaper than, the chat model.
ChunkA piece of a document, sized so it can be retrieved on its own.
Vector database / vector searchStorage that finds items with similar embeddings. Some regular databases offer it too — Postgres, for example, through the pgvector extension.
Top-kHow many chunks are retrieved per question.
Hybrid searchCombining meaning-based search with ordinary keyword search, which helps with exact terms like product codes.
Context windowThe maximum amount of text a model can take in one request.
Grounding / citationsTying the answer to specific retrieved sources and showing them.

Common misconceptions

  • "RAG trains the model on my data." It doesn't. Nothing about the model changes; it just reads extra text at question time.
  • "RAG means no more hallucinations." It means fewer. If retrieval misses the right chunk, or the answer isn't in your documents, the model may still guess. Instruct it to say "I don't know" and test with questions whose answers you know.
  • "I need a special vector database." Often you don't. For modest amounts of data, a vector extension on the database you already have, or even plain keyword search, can be enough.
  • "Bigger chunks are better." Too big and you send lots of irrelevant text (more cost, more confusion); too small and a chunk loses the context that made it meaningful. Expect to adjust.
  • "Just upload everything." Outdated, duplicated or contradictory documents produce contradictory answers. Cleaning the source material is often the biggest improvement you can make.
  • "Anyone can see any document." Only if you build it that way. If some documents are private to certain users, the search itself must filter by who is asking — otherwise the assistant can quote one customer's data to another.

What it costs

A RAG feature usually has three running costs: embedding your documents (small, and mostly one-off), embedding each question (tiny), and the language model call (the main cost, which grows with how many chunks you send). Plus storage in your database. Your AI provider bills these to your own account, so set a spending limit there.

What to ask your AI builder

  • "Before building RAG, check: is our content small enough to just include in the prompt?"
  • "Split the help articles into chunks of a few paragraphs, keep each chunk's title and source link, and store embeddings in our database."
  • "Use hybrid search — meaning-based plus keyword — so product codes still match."
  • "Tell the model to answer only from the retrieved sources, say when it doesn't know, and show the sources under each answer."
  • "Only search documents the signed-in user is allowed to see."
  • "Re-embed a document automatically whenever it's edited."
  • "Keep the AI provider key on the server in a secret, never in the browser."
  • "Log each question, the chunks retrieved and the answer so I can review bad answers."

RAG in Mythex

Mythex doesn't include a ready-made knowledge-base or RAG product. What it does is write the code for one inside your app: you bring your documents and your own AI provider key, stored as a project secret, and the agent builds the retrieval and chat parts, using the project's Postgres database for storage. Provider usage bills your provider account, not Mythex credits. The docs suggest starting with one simple AI feature and adding RAG once that works — see building an AI app, and our guides on building an AI chatbot app and adding AI features.

Questions

What is RAG in simple terms?

RAG (retrieval-augmented generation) is a way to make an AI model answer from your own information. Before the model replies, your app searches your documents for the passages most relevant to the question and includes them in the prompt, so the answer is based on that text rather than only on what the model learned in training.

Does RAG train the AI on my data?

No. RAG doesn't change the model at all. It looks up relevant text at the moment a question is asked and passes it to the model along with the question. Update a document and the next answer can use the new version straight away.

Is RAG the same as fine-tuning?

No. Fine-tuning changes a model's behaviour by training it further on examples, which is slower and better for style or format. RAG gives the model facts to read at question time, which is better for knowledge that changes or must be quoted accurately.

Does RAG stop AI hallucinations?

It reduces them but doesn't eliminate them. If the search finds the wrong passages, or the answer isn't in your documents, the model can still guess. Good RAG apps tell the model to say when it doesn't know and show the sources they used.

Keep reading

  • Frontend vs Backend: What's the Difference? — The frontend is what users see in the browser; the backend runs on a server and handles data, logic and security. How the two fit together, with an example.
  • How Domains and DNS Work: A Guide for Non-Developers — How domain names and DNS connect example.com to your app: registrars, nameservers, A, CNAME, MX and TXT records, propagation, and connecting a custom domain.
  • How to Add AI Features to Your App — Add summaries, chat, data extraction and classification to your app with an LLM API — keeping keys safe, costs under control and output trustworthy.
  • How to Build an AI Chatbot App: Keys, Prompts, Costs and Safety — How to build your own AI chatbot app: choosing an LLM provider, keeping API keys safe, writing the system prompt, controlling costs and adding guardrails.
  • How to Use LLM APIs: Tokens, Costs, Keys and Your First AI Feature — What an LLM API is, how tokens, context windows and per-token pricing work, how to keep your API key safe, and how to add a first AI feature to your app.
  • Native Apps vs Progressive Web Apps: Which Do You Need? — Native apps vs progressive web apps (PWAs): what each can do, iPhone limits as of September 2026, costs, and how to choose for your first version.

Start building free · Templates · Docs