RAG feeds relevant documents into the prompt so the model can answer from them.
publishedragretrievaloverviewUpdated 2026-10-05
Retrieval-Augmented Generation
A base model only knows what was in its training data up to its knowledge cutoff. Retrieval-augmented generation (RAG) is the standard way to fix that: at query time, find the documents relevant to the user's question and place them in the model's context so it can answer from them rather than from memory. RAG is how you make a model answer questions about your own, current, or private data without retraining it.
This section walks the full pipeline, one stage at a time.
Documents in this section
- RAG End to End — the whole pipeline and how the stages fit together. Start here.
- Chunking Strategies — splitting documents into retrievable units, and why chunk size is a real design decision.
- Vector Databases — storing and searching embeddings at scale.
- Retrieval and Reranking — getting from "roughly relevant" to "actually the best passages."
- Evaluating RAG — measuring whether your pipeline actually retrieves the right thing and answers faithfully.
Prerequisite: read Embeddings first — the whole section assumes you know what a vector is and what cosine similarity means.