Retrieval Augmented Generation (RAG)
1 · The lesson
readAn LLM does not know your company's documentation, your customer's order history, the contents of the PDF a user just uploaded, or anything published after its training cutoff. RAG — Retrieval Augmented Generation — is how you bridge that gap without retraining the model. You search a knowledge base for the chunks relevant to the user's question, paste them into the prompt, and ask the LLM to answer using those chunks.
The pattern is simple. Doing it well — at production quality, latency, and cost — is one of the hardest things in applied LLM work. This lesson is the working architecture, the failure modes, and the smallest-possible RAG you can build in fifty lines of Python.
Run locally with
pip install anthropic sentence-transformers numpyandANTHROPIC_API_KEYset. Expected outputs shown in comments.sentence-transformersdownloads a model on first run (~80 MB).
1. The Problem RAG Solves
An LLM is a frozen statistical model. It cannot:
- Know facts more recent than its training cutoff
- Know your facts — internal docs, user data, proprietary content
- Reliably cite sources for what it says
Three workarounds exist:
| Approach | When to use |
|---|---|
| Long context — paste the whole knowledge base into the prompt | Small corpora (< 100K tokens), one-shot |
| Fine-tuning | Style/format adaptation, not new knowledge — see next lesson |
| RAG | Anything bigger; updates over time; need citations |
RAG wins for almost every real "the LLM needs to know X" problem. It is updatable (just re-index when the docs change), cheap (no training), and citable (you know exactly which chunks were used).
2. The Pipeline — Five Steps
RAG, in five stages. Every production system follows this shape; the differences are in how each stage is implemented.
flowchart LR
A[Documents] --> B[Chunk]
B --> C[Embed]
C --> D[(Vector DB)]
Q[User query] --> E[Embed query]
E --> F[Retrieve top-k]
D --> F
F --> G[Build prompt]
G --> H[LLM]
H --> R[Answer + citations]1. Chunk documents into passages small enough to embed meaningfully and big enough to carry context.
2. Embed each chunk — turn text into a fixed-size vector that captures meaning.
3. Index the vectors in a database that supports fast nearest-neighbour search.
4. At query time: embed the question, retrieve the top-k nearest chunks.
5. Build the prompt with the retrieved chunks as context, send to the LLM.
Each stage has its own failure modes. Bad chunks → no relevant text retrievable. Bad embeddings → wrong chunks retrieved. Bad prompt → model ignores the context. We will hit all three.
3. Embeddings — Text to Vectors
An embedding model maps a piece of text to a fixed-length vector (typically 384, 768, 1024, or 1536 dimensions). The geometry is meaningful: texts about similar topics land near each other, irrelevant texts land far apart.
Three families of embeddings, roughly in order of cost:
| Source | Examples | Use when |
|---|---|---|
| Open, local | sentence-transformers (BGE, E5, all-MiniLM) | Privacy, cost, offline use |
| API-hosted | OpenAI text-embedding-3-*, Voyage, Cohere | Best quality, no infrastructure |
| Built into the LLM provider | Anthropic via Voyage, Google Gecko | Co-located with your inference |
The standard open-source workhorse is sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("BAAI/bge-small-en-v1.5") # 384-dim, ~30 MB vec = model.encode("Python is a programming language.") print(vec.shape) # (384,)
Run a sentence through model.encode() and you get back a numpy array. Run a thousand sentences in a list and you get a (1000, 384) matrix. That is the basic operation.
Similarity is cosine similarity — the dot product of the two unit-length vectors. Higher = more similar.
import numpy as np def cosine(a, b): return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
Most embedding models output normalised vectors (unit length), so cosine reduces to a plain dot product. Faster, same answer.
4. Chunking — The First Real Problem
You cannot embed an entire 50-page PDF as one vector — too much information collapses into 384 numbers and similarity becomes meaningless. So you split into chunks.
| Strategy | How | Trade-off |
|---|---|---|
| Fixed-size | Every N characters or tokens | Simple, but cuts mid-sentence and breaks coherent ideas |
| Sentence / paragraph | Split on . or \n\n | Respects structure; varying chunk sizes |
| Semantic | Split where embedding similarity between adjacent sentences drops | Best topical coherence; slow to compute |
| Hierarchical | Multi-level: indexed at small chunk, retrieved with surrounding parent paragraph | Best of both worlds; more complex |
| Token-aware with overlap | ~500 tokens, 50-100 token overlap between adjacent chunks | The pragmatic default |
The overlap trick deserves emphasis: if a relevant sentence falls exactly at a chunk boundary, neither chunk has the full context. A 10-20% overlap between adjacent chunks fixes this. Cost: slightly more chunks to embed and store.
For structured documents (markdown, code), split on natural boundaries (headings, function definitions) rather than character counts. The structure carries semantic meaning that fixed-size chunking destroys.
5. Vector Databases — A Quick Tour
Once you have thousands or millions of embedding vectors, you need fast nearest-neighbour search. Options range from "a numpy array" to "managed cloud service":
| Tool | When to use |
|---|---|
| In-memory numpy / scikit-learn | < 10K chunks, prototyping |
| FAISS (Meta) | In-memory or memory-mapped, fast, library only — no server |
| Chroma | File-based, embedded, "SQLite of vector DBs". Great for small-to-medium apps |
| Qdrant / Weaviate / Milvus | Self-hosted servers, production scale, rich filtering |
| Pinecone / Vespa / Turbopuffer | Managed cloud services, no ops |
| Postgres + pgvector | You already have Postgres; modest scale |
For most production apps starting out, Chroma or Qdrant (run locally in Docker) is plenty. You typically don't need managed cloud until you're at hundreds of millions of vectors or need cross-region replication.
All of them implement Approximate Nearest Neighbour (ANN) algorithms — HNSW, IVF, ScaNN — that trade a tiny amount of recall for huge speed-ups vs exact search. Exact cosine over 10M vectors is ~seconds; ANN is ~milliseconds.
6. A Minimum-Viable RAG — 40 Lines
To build intuition, here is the smallest plausible RAG with no vector DB, just numpy:
# Run locally with `pip install sentence-transformers anthropic numpy`. import numpy as np from sentence_transformers import SentenceTransformer from anthropic import Anthropic # 1. Tiny knowledge base (5 chunks) CHUNKS = [ "Python is a high-level, interpreted programming language created by Guido van Rossum in 1991.", "Python's official package manager is pip. The Python Package Index (PyPI) hosts over 500,000 packages.", "Decorators in Python are functions that modify other functions, denoted with the @ symbol.", "Python uses indentation, not braces, to define blocks of code. Standard is 4 spaces.", "The Global Interpreter Lock (GIL) means CPython only runs one thread of Python bytecode at a time.", ] # 2. Embed the chunks once embedder = SentenceTransformer("BAAI/bge-small-en-v1.5") chunk_vecs = embedder.encode(CHUNKS, normalize_embeddings=True) # (5, 384) def retrieve(query: str, k: int = 3) -> list[str]: """Return the top-k most similar chunks to the query.""" q = embedder.encode(query, normalize_embeddings=True) # (384,) scores = chunk_vecs @ q # cosine = dot, since normalised top_idx = np.argsort(scores)[::-1][:k] return [CHUNKS[i] for i in top_idx] def answer(query: str) -> str: context = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(retrieve(query))) prompt = ( f"Answer the question using ONLY the context below. If the answer " f"is not in the context, say 'I don't know'. Cite chunks by their [N].\n\n" f"<context>\n{context}\n</context>\n\n" f"<question>{query}</question>" ) client = Anthropic() resp = client.messages.create( model="claude-opus-4-7", max_tokens=300, messages=[{"role": "user", "content": prompt}], ) return resp.content[0].text print(answer("Who created Python and when?")) # Python was created by Guido van Rossum in 1991. [1]
Forty lines. Real retrieval, real LLM, real citations. Swap CHUNKS for a chunked PDF and you have something you could ship. The rest of this lesson is about making it good.
7. Retrieval Quality Is Where You Live or Die
Here is the central truth: if you retrieve the wrong chunks, the LLM cannot recover. No amount of prompt engineering fixes "the relevant passage was never in the context". RAG quality is dominated by retrieval quality, which means embedding quality, chunking strategy, and ranking.
The single biggest improvement most teams make to a working-but-mediocre RAG: better retrieval, not better generation.
7.1 Hybrid Retrieval — Vector + Lexical
Dense embeddings are great at semantic similarity but bad at proper nouns, error codes, and exact-string matches. "Error E_INVALID_TOKEN" might not embed near a chunk that contains that exact string. Mix in BM25 (a 1990s lexical relevance score) and you get the best of both worlds:
from rank_bm25 import BM25Okapi # At index time tokenised = [c.lower().split() for c in CHUNKS] bm25 = BM25Okapi(tokenised) # At query time def hybrid_retrieve(query: str, k: int = 5, alpha: float = 0.5): q_vec = embedder.encode(query, normalize_embeddings=True) vec_scores = chunk_vecs @ q_vec # in [-1, 1] bm25_scores = bm25.get_scores(query.lower().split()) bm25_norm = bm25_scores / (bm25_scores.max() + 1e-9) # crude normalisation combined = alpha * vec_scores + (1 - alpha) * bm25_norm top_idx = np.argsort(combined)[::-1][:k] return [CHUNKS[i] for i in top_idx]
setup added so this can run · defines CHUNKS, chunk_vecs, embedder, np
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) CHUNKS = _AutoMock('CHUNKS') chunk_vecs = _AutoMock('chunk_vecs') embedder = _AutoMock('embedder') np = _AutoMock('np')
alpha=0.5 is a sensible starting point. Tune on your eval set.
7.2 Re-Ranking — Retrieve Many, Keep Few
A second-stage re-ranker dramatically improves quality on the same retrieval budget. The pattern:
1. First stage (cheap, fast): retrieve top 20-50 candidates with vector / BM25.
2. Second stage (expensive, accurate): score each candidate with a cross-encoder that takes (query, chunk) as a pair and outputs a relevance score.
3. Keep the top 3-5 after re-ranking, send to the LLM.
from sentence_transformers import CrossEncoder reranker = CrossEncoder("BAAI/bge-reranker-base") # ~280 MB def rerank(query: str, candidates: list[str], k: int = 5): pairs = [[query, c] for c in candidates] scores = reranker.predict(pairs) ranked = sorted(zip(candidates, scores), key=lambda x: x[1], reverse=True) return [c for c, _ in ranked[:k]]
Why it works: a bi-encoder (embedding model) has to encode query and chunk independently — neither sees the other. A cross-encoder runs them together through one transformer and can produce a much sharper relevance judgement. Slower per-pair, but you only run it on 50 candidates, not the whole corpus.
This pattern — fast first stage, slow accurate second stage — is the workhorse of every search system since 2010s information retrieval. RAG inherited it for good reason.
8. The RAG Prompt Template
A workable template that handles the common cases:
Use ONLY the context below to answer the question. If the answer is not in the context, say "I don't have that information in the provided sources". Cite the source chunks by their [N] identifier wherever you make a factual claim. <context> [1] {chunk_1} [2] {chunk_2} [3] {chunk_3} </context> <question> {user_question} </question>
The key elements:
- "ONLY" — actively discourages drawing on training data, which is where hallucinations hide.
- "If not in the context, say..." — gives the model an explicit fallback. Without this, it makes something up.
- Cite by [N] — produces verifiable output. You can render the citations as links to the source documents.
- XML tags — keep
contextandquestioncleanly separable so the model does not confuse them.
Always include chunk IDs. Always require citations. Always render the citations to the user (or at least surface them at debug time). A citation that traces back to the source document is the difference between "AI said X" and "AI said X, here is exactly where it came from in your docs".
9. Evaluating RAG — Three Metrics
Generic LLM eval (does the answer look good?) is not enough for RAG. You need to separate retrieval failures from generation failures:
| Metric | Measures | How |
|---|---|---|
| Retrieval recall@k | Did we retrieve a chunk that contains the answer? | Hand-label a chunk as "ground truth" for each test query; check if it's in the top-k |
| Answer faithfulness | Does the answer follow from the retrieved chunks, or is it hallucinated? | LLM-as-judge: "Is every claim in the answer supported by the context?" |
| Answer relevance | Does the answer address the question? | LLM-as-judge: "Does the answer actually answer the question?" |
A retrieval-recall@5 of, say, 0.7 means 30% of queries are unanswerable by the LLM regardless of how brilliant the prompt is. Fix retrieval first.
Frameworks like Ragas and TruLens automate these metrics. Worth setting up before you have lots of users.
10. The Lost-in-the-Middle Problem
Long context windows don't make context use uniform. Liu et al. (2023) showed that information in the middle of a long context is recalled significantly less reliably than information at the beginning or end. A relevant chunk placed in the middle of a 50-chunk context might as well not be there.
Implications for RAG:
- Keep retrieved context tight. 5 great chunks beat 50 mediocre ones.
- Order matters. Put the most relevant chunks at the start or end of the context block, not in the middle.
- Re-ranking helps twice — better top-k, and the top-1 chunk ends up at a strong position.
Modern frontier models have improved on this, but the effect has not disappeared. Don't dump everything you have into context and assume the model will sort it out.
Common Mistakes
1. Chunks too large or too small
500-2000 character chunks (~100-400 tokens) are usually the sweet spot. Chunks of 50 characters lose context; chunks of 10,000 characters water down embedding similarity (one vector cannot represent ten different ideas well).
2. Not deduplicating retrieved chunks
If your corpus has near-duplicates (mirrored docs, identical FAQs across products), the top-5 might all be variants of the same passage. The LLM sees one piece of information five times and nothing else. Run a similarity check on retrieved chunks and drop near-duplicates before building the prompt.
3. Skipping chunk overlap
Without overlap, key sentences at chunk boundaries are split across two chunks, and neither has the full context. A 10-20% overlap is cheap and noticeably improves recall.
4. Not testing retrieval separately from generation
"My RAG is bad" almost always means one of: retrieval is bad, prompting is bad, or both. Build a retrieval eval set first (questions + ground-truth chunks); fix recall@k before tuning prompts.
5. Thinking RAG solves hallucination
RAG constrains hallucination by giving the model facts to lean on. It does not eliminate it. The model can still:
- Misread a chunk and produce a wrong answer
- Combine information from multiple chunks incorrectly
- Add plausible-sounding details not in any chunk
- Cite the wrong chunk for a correct fact
Verification (LLM-as-judge, human review on samples, citation-checking automation) is still needed.
6. Embedding the wrong text
Embedding a chunk of source code and querying with English may underperform — the embedding model was trained on natural language, not code. Use a code-specialised embedding model (microsoft/codebert-base, voyage-code-3) or embed natural-language descriptions alongside the code.
🎯 Your Turn — Build a Five-Chunk RAG
Write rag_answer(query, chunks) that:
- takes a list of text chunks and a query,
- embeds the chunks (use
sentence-transformersmodel"BAAI/bge-small-en-v1.5"), - embeds the query,
- returns the top-3 chunks ranked by cosine similarity (don't call the LLM — focus on the retrieval part),
- returns them as a list of
(score, chunk)tuples, highest score first.
Use the provided five-chunk knowledge base. Verify on a few queries.
# Run locally with `pip install sentence-transformers numpy`. import numpy as np from sentence_transformers import SentenceTransformer CHUNKS = [ "The capital of France is Paris, known for the Eiffel Tower.", "Python is a programming language created by Guido van Rossum.", "Mount Everest is the tallest mountain in the world, at 8,849 metres.", "The Pacific Ocean is the largest ocean, covering about 30% of Earth's surface.", "JavaScript runs in every browser and is one of the most popular languages.", ] def rag_answer(query: str, chunks: list[str], k: int = 3) -> list[tuple[float, str]]: # TODO 1: load the embedding model "BAAI/bge-small-en-v1.5" # TODO 2: embed chunks with normalize_embeddings=True -> matrix of shape (N, 384) # TODO 3: embed query with normalize_embeddings=True -> vector of shape (384,) # TODO 4: scores = matrix @ vector (since normalised, this is cosine similarity) # TODO 5: get top-k indices (descending) # TODO 6: return [(float(scores[i]), chunks[i]) for i in top_idx] ... if __name__ == "__main__": for score, chunk in rag_answer("Tell me about programming languages", CHUNKS): print(f"{score:.3f} {chunk}")
Hint 1 — Cosine via dot product
When both vectors are unit-length (whichnormalize_embeddings=True ensures), cosine similarity equals the plain dot product. So chunk_vecs @ q_vec gives you all similarities at once as a 1-D array — no need for a loop or explicit cosine calculation.
Hint 2 — Top-k indices
np.argsort(scores) sorts ascending. Reverse with [::-1], then take [:k] for the top k. Or use np.argpartition if you care about speed on huge arrays (you don't, here).
Show full solution
# Run locally with `pip install sentence-transformers numpy`. import numpy as np from sentence_transformers import SentenceTransformer CHUNKS = [ "The capital of France is Paris, known for the Eiffel Tower.", "Python is a programming language created by Guido van Rossum.", "Mount Everest is the tallest mountain in the world, at 8,849 metres.", "The Pacific Ocean is the largest ocean, covering about 30% of Earth's surface.", "JavaScript runs in every browser and is one of the most popular languages.", ] # Load once at module level — the model takes ~1s to load _embedder = SentenceTransformer("BAAI/bge-small-en-v1.5") def rag_answer(query: str, chunks: list[str], k: int = 3) -> list[tuple[float, str]]: """Return the top-k chunks ranked by cosine similarity to the query.""" chunk_vecs = _embedder.encode(chunks, normalize_embeddings=True) # (N, 384) q_vec = _embedder.encode(query, normalize_embeddings=True) # (384,) scores = chunk_vecs @ q_vec # (N,) top_idx = np.argsort(scores)[::-1][:k] return [(float(scores[i]), chunks[i]) for i in top_idx] if __name__ == "__main__": print("Query: 'Tell me about programming languages'") for score, chunk in rag_answer("Tell me about programming languages", CHUNKS): print(f" {score:.3f} {chunk}") print("\nQuery: 'How high is Everest?'") for score, chunk in rag_answer("How high is Everest?", CHUNKS): print(f" {score:.3f} {chunk}") # Example output: # Query: 'Tell me about programming languages' # 0.789 JavaScript runs in every browser and is one of the most popular languages. # 0.762 Python is a programming language created by Guido van Rossum. # 0.421 The Pacific Ocean is the largest ocean, covering about 30% of Earth's surface. # # Query: 'How high is Everest?' # 0.881 Mount Everest is the tallest mountain in the world, at 8,849 metres. # 0.376 The Pacific Ocean is the largest ocean, covering about 30% of Earth's surface. # 0.339 The capital of France is Paris, known for the Eiffel Tower.
What this gets right:
- Re-embedding chunks every call is fine for 5 chunks but wasteful in production. Cache
chunk_vecsoutside the function or persist it (FAISS / Chroma) for larger corpora. normalize_embeddings=Truemeans dot product equals cosine similarity — no separate normalisation step needed.@operator does matrix-vector multiplication; numpy vectorises the whole similarity computation.float(...)cast converts numpy scalars to plain Python floats so the tuple prints cleanly.
The chasm between this and production:
- Persisted vector store — for thousands+ chunks, use Chroma, Qdrant, or FAISS so embeddings don't recompute on every query.
- Hybrid retrieval — combine with BM25 for queries with proper nouns or rare strings.
- Re-ranking — cross-encoder on top-20 → keep top-3. Significantly better quality.
- Chunking strategy — your real chunks come from documents; you need to split intelligently with overlap.
- Evaluation — measure recall@k on a labelled test set; iterate.
- Pass retrieved chunks into an LLM — this exercise stops at retrieval to keep focus; the next step is the prompt-and-call pattern from Section 6.
The shape — embed corpus once, embed query, dot product, top-k — is the kernel of every RAG system from a tutorial to production-scale search. Everything else is engineering around it.
What You Learned
- RAG bridges the gap between an LLM's frozen knowledge and your data: retrieve relevant chunks, paste into the prompt, generate.
- The five-stage pipeline: chunk → embed → index → retrieve → prompt.
- Embeddings map text to vectors; cosine similarity (or dot product on normalised vectors) measures relevance.
- Chunking is its own decision — fixed-size, paragraph, semantic, hierarchical, with overlap. Get this wrong, and nothing else helps.
- Vector databases range from in-memory numpy to managed cloud — pick the simplest that fits your scale.
- Retrieval quality dominates RAG quality. Hybrid retrieval (vector + BM25), re-ranking with a cross-encoder, and tight top-k all matter.
- A working RAG prompt uses XML tags, says "ONLY use this context", offers a fallback, and demands citations.
- Evaluate retrieval separately from generation: recall@k, faithfulness, relevance.
- Lost-in-the-middle is real: keep context tight, put the best chunks at the edges.
- RAG constrains hallucination; it does not eliminate it. Verification still matters.
Next: Fine-tuning & Training LLMs — when prompts and RAG aren't enough, when fine-tuning is the right tool, and when (usually) it is not.