What Is RAG? A Guide for Internal Documentation
What is RAG? It's how an AI answers questions from your own docs instead of guessing. Here's how it works, with a real example inside a self-hosted wiki.
TL;DR
- RAG (Retrieval-Augmented Generation) lets an AI answer questions using your own documentation, instead of guessing from what it learned during training.
- It works in two steps: it retrieves the most relevant pages, then generates an answer from them.
- Because the AI reads your real pages before it answers, it's less likely to make things up, and it stays current when you update your docs.
- Docmost's AI Answers feature is a working example of RAG built into a wiki, and this guide walks through how it's set up.
Keyword search finds pages by the words you type, so it works best when you already know the words the author used. In a wiki with hundreds of pages, a search for "refund approval" can miss a page titled "Escalation process for billing exceptions," even if that page contains the answer.
RAG is one of the main ways AI tools fix this problem. It finds pages by what they're about, not by matching words, and then has an AI write an answer from what it found. This guide covers what RAG is, how it works, and how it runs inside a real wiki, so you can judge whether an AI search feature will help your team find answers.
RAG Key Facts
|
Aspect |
What it means |
|
What it stands for |
Retrieval-Augmented Generation |
|
What it solves |
Lets an AI answer from your team's current documentation, not just what it learned during training |
|
The two important steps |
1) Retrieve the most relevant pages; 2) Generate an answer from them |
|
What it needs |
Your content stored so it can be searched by what it's about (a vector database), and an AI model that can write the answer |
|
Why it matters for internal docs |
Answers come from real pages and get updated as the wiki changes |
|
Docmost example |
AI Answers (paid edition), which searches your workspace using vector embeddings and runs on an OpenAI, Google Gemini, or self-hosted Ollama model that your team connects |
What Does RAG Mean?
RAG stands for Retrieval-Augmented Generation. In AI, it's a way of making a large language model (LLM) answer questions using "information" it looks up at the moment you ask, instead of relying only on what it learned during training. The "retrieval" part finds the relevant content, and the "generation" part is the AI writing an answer from it.
For internal documentation, that "information" is your own wiki: your policies, runbooks, onboarding guides, and project notes. You'll sometimes see the whole setup called a RAG model or a RAG system. Both mean the same thing, which is an AI model paired with a search step that feeds it the right material before it answers.
How RAG Works
People often call the process below a RAG pipeline. Some of the work happens ahead of time, when your documentation gets prepared, and the rest runs every time someone asks a question.
Step 1: Retrieval (and the Prep Behind It)
Before anyone asks anything, the system prepares your documentation. It splits pages into smaller pieces and turns each piece into an embedding, which is a list of numbers that represents what the text means. Those embeddings are stored in a vector database, a type of database built to find entries that are close to each other in meaning.
When a question gets asked, it's turned into an embedding too, and the system looks for the pieces that are closest to it in meaning. That's how a search for "refund approval" can find a page about "billing exceptions." The words don't match, but the meaning does.
Step 2: Generation
Next, the AI model receives the retrieved pieces along with the original question and writes an answer based on that material. Since the answer is built from content it just pulled from your docs, the model doesn't have to guess as much. This is the main reason RAG means fewer made-up answers, which people in AI call hallucinations.
RAG vs. LLM: What's the Difference?
An LLM is the model that writes text, while RAG is a setup that gives that model the right information before it writes. So when people compare the two, they're comparing an LLM working from memory with an LLM that gets to read your documents first.
|
Aspect |
LLM on its own |
LLM with RAG |
|
Where answers come from |
Only what it learned in training |
What it learned in training, plus your current documents |
|
Keeps up when your docs change |
No, the model would need retraining |
Yes, because retrieval happens when the question is asked |
|
Chance of hallucinated answers |
Higher |
Lower, since answers are gotten from retrieved content |
Your documentation changes constantly, and retraining a model every time someone edits a page isn't practical. That's why AI search tools built on top of your documentation tend to use retrieval.
An Example Of How RAG Works Inside Docmost
Docmost's AI Answers feature is RAG built directly into the wiki, so you can see retrieval and generation running in a real product. It's part of Docmost's paid edition.
- Your server gets connected to an AI provider: Whoever runs your Docmost server sets up the provider: OpenAI, Google Gemini, or a local model running through Ollama. They also add pgvector, an extension that lets Postgres, the database Docmost runs on, store and search vector embeddings.
- An admin turns on AI Answers: A workspace admin switches on AI-powered search under Settings > AI settings, and Docmost starts generating embeddings for your existing pages in the background.
- New and edited pages keep up: When someone creates or updates a page, Docmost generates fresh embeddings for it in the background, so answers keep up with changes shortly after they're made.
- Someone asks a question: A user opens search, switches on the AI Answers toggle, and types something like "Who approves refunds above the usual limit?" They get a written answer built from the workspace's own pages.
- Answers respect permissions: AI Answers only searches the spaces and pages that person has access to, so an answer can't come from a space they aren't allowed to see.
That last point is a big deal. HR and finance spaces in a company wiki are often restricted, and an AI search feature that ignored those limits would be a problem no matter how good its answers were. If you're looking at this kind of feature across tools, see how self-hosted wikis with AI built in compare.
Why RAG Is Important for Internal Documentation
- People find answers without guessing the right keywords: Retrieval works on meaning, so a new hire doesn't need to know what the author called the page before they can find it.
- Answers stay tied to what your team wrote: Because the AI works from retrieved content, its answers come from your documentation instead of whatever the model assumes.
- You decide where your data goes: With a self-hosted wiki, your documentation never leaves infrastructure you control. If your team also runs both the embedding model and the answering model through Ollama on its own servers, the questions and answers stay on that infrastructure too, instead of going out to a cloud AI provider.
One thing RAG can't fix is documentation that's wrong or out of date. It retrieves whatever is closest in meaning, so if two pages disagree, the answer may come from the wrong one. Keeping pages current and cleaning up duplicates do as much for AI search as the model you choose. If your team's documentation currently lives in another tool, the import and export guide shows how to bring your existing pages into Docmost.
Frequently Asked Questions
- What Does RAG Mean?
RAG stands for Retrieval-Augmented Generation. It's a way to make an AI answer questions using content it looks up from a real source, such as your company wiki, instead of relying only on what it learned during training.
- What Is an Example of RAG?
Docmost's AI Answers feature is one. It turns your workspace pages into embeddings, finds the ones that match a question in meaning, and writes an answer from them. A support chatbot that answers from a company's help center works the same way.
- What's the Difference Between an LLM and RAG?
An LLM is the model that generates text. RAG is a setup that retrieves relevant content and hands it to the LLM before it answers. The two work together, with RAG adding a lookup step so the model has something current to work from.
- Is ChatGPT a RAG LLM?
Not on its own. ChatGPT's underlying model answers from what it learned in training. Some ChatGPT features add a retrieval step on top, such as searching the web or looking through large files you upload, and in those cases it's working in a RAG-style way.
- How Is RAG Different From Fine-Tuning?
Fine-tuning retrains the model itself on new data. RAG leaves the model alone and gives it fresh material at question time, so updating your docs updates the answers without retraining anything.
- Does RAG Stop AI Hallucinations?
It reduces them but doesn't remove them completely. The model can still misread a retrieved passage, and when nothing relevant exists in your docs, a well-built system should say so instead of guessing.
- Do I Need a Vector Database to Use RAG With a Wiki?
Yes, some form of one. Your content gets converted into embeddings and stored somewhere that can be searched by meaning. In Docmost, that job is handled by pgvector, which your team adds to Postgres before turning on AI search.
Bring AI Search to Your Own Documentation
Docmost's AI Answers brings RAG into a wiki you can self-host, with your choice of AI provider. Contact sales to talk through the paid edition and how AI search would work for your team.