# RAG (Retrieval-Augmented Generation)

URL: https://softwaredictionary.org/terms/rag
Category: AI & Machine Learning
Last updated: 2026-09-29
Pronunciation: RAG

In short: RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.

## What is RAG?

Retrieval-Augmented Generation, or RAG, combines a search step with a language model. When a user asks a question, the system first retrieves the most relevant pieces of information from a knowledge source, such as company documents or a help center, and adds them to the prompt. The LLM then generates its answer from that supplied context instead of relying only on what it learned during training.

A typical RAG pipeline has two phases. Ahead of time, documents are split into smaller chunks, each chunk is turned into an embedding, and the embeddings are stored in a vector index. At question time, the question is embedded too, the closest chunks are retrieved, and they are inserted into the prompt together with an instruction such as 'answer using only the sources below'.

An everyday analogy is an open-book exam: instead of answering from memory, the student looks up the right pages first. RAG is widely used for internal knowledge assistants, customer support bots, and documentation search, because it reduces hallucinations and lets answers point to their sources.

RAG is often compared with fine-tuning. Fine-tuning trains the model further on your data, which changes its behavior or style but is slow and costly to update, while RAG leaves the model unchanged and simply gives it fresh information on every request. For knowledge that changes often, RAG is usually the simpler and cheaper choice.

## Key takeaways

- RAG retrieves relevant data first, then asks the LLM to answer using it.
- It usually relies on embeddings and a vector index to find relevant text chunks.
- It keeps answers current without retraining the model.
- It reduces, but does not eliminate, hallucinations.
- The quality of retrieval largely determines the quality of the answer.

## Example: A minimal RAG flow

```typescript
// embed, vectorIndex and llm are placeholders for real services
async function answer(question: string): Promise<string> {
  // 1. Retrieve: find the text chunks most similar to the question
  const queryVector = await embed(question);
  const chunks = await vectorIndex.search(queryVector, { topK: 3 });

  // 2. Augment: add the retrieved text to the prompt
  const sources = chunks.map((chunk) => chunk.text).join("\n\n");
  const prompt = `Answer using only these sources:\n${sources}\n\nQuestion: ${question}`;

  // 3. Generate: the LLM writes an answer grounded in the sources
  return llm.generate(prompt);
}
```

## Frequently asked questions

**What is the difference between RAG and fine-tuning?**

RAG gives the model relevant information at request time without changing the model, while fine-tuning trains the model further on your own data. RAG is better for facts that change often; fine-tuning is better for teaching a consistent style, format, or specialized behavior.

**Does RAG stop hallucinations?**

RAG reduces hallucinations by giving the model real sources to work from, but it does not eliminate them. The model can still misread the sources, and if retrieval returns the wrong documents, the answer can still be wrong.

**Do I need a vector database for RAG?**

Not necessarily. Many RAG systems use a vector database, but a regular database with vector search support, or even classic keyword search, can also serve as the retrieval step.

## Sources

- [Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020)](https://arxiv.org/abs/2005.11401)

---

Software Dictionary: https://softwaredictionary.org/ · https://softwaredictionary.org/llms.txt
