Level 9 — Retrieval-Augmented Generation
Level 8 changed model behavior by changing model parameters. Level 9 keeps the model fixed and instead changes which external evidence is supplied at request time. That pattern is called retrieval-augmented generation, or RAG.
Learning goal
Build a traceable evidence pipeline:
documents
→ chunks + metadata
→ retrieval representation
→ candidate search
→ hybrid search/reranking
→ query rewrite when needed
→ grounded prompt
→ answer + citations
→ retrieval and answer evaluation
→ freshness/access/update controls
What mastery looks like
By the end of the level, you should be able to:
- explain when RAG is more appropriate than fine-tuning;
- split documents into chunks without losing source identity;
- attach metadata needed for citation, filtering, freshness, and access control;
- explain retrieval embeddings separately from generation hidden states;
- calculate dot product and cosine similarity;
- build a small deterministic vector index;
- compare lexical, vector, and hybrid retrieval;
- separate candidate retrieval from reranking;
- rewrite queries without silently changing user intent;
- construct prompts that keep trusted instructions separate from retrieved evidence;
- validate that citations point to supplied sources;
- measure retrieval success separately from answer quality;
- diagnose whether a failure began in indexing, retrieval, reranking, prompting, or generation;
- enforce freshness and access constraints outside prompt wording;
- package an evidence-backed RAG workflow with reproducible source and evaluation identities.
Mini checkpoint
Complete the checkpoint after L9.7. Before moving to grounded generation, you should be able to trace a query through chunks, embeddings/scores, retrieval, hybrid search, and reranking.
Level project
After L9.14, complete Evidence-Backed RAG. The default local path is deterministic and does not require a hosted embedding or LLM provider. L9.14 includes an optional real sentence-embedding exercise, with an additional real generator call available, while preserving the same retrieval, grounding, and evaluation checks.
RAG changes the evidence path, not the model's stored knowledge
A retrieval-augmented system answers a question in two broad phases. First it decides which external records are eligible and relevant. Then it gives selected evidence to a generator. Those phases can fail independently. A correct document may exist but never be retrieved, or the right chunk may reach the prompt and still be misread by the generator.
That separation is the central idea of this level. You will keep source identity, chunk identity, retrieval score, rank, supplied context, and final citations visible. When an answer is wrong, do not begin with “How do I rewrite the prompt?” Begin with a different question: “At which boundary did the relevant evidence disappear or become corrupted?”
Follow one question through the pipeline
Imagine a product manual containing the sentence “XR-4172 warranty: 18 months.” The source document first receives stable metadata such as product, version, tenant, and update time. Chunking creates smaller retrieval units while preserving a link back to the source. An embedding or lexical representation turns the query and chunks into comparable signals. Search produces candidates, an optional reranker changes their order, and access/freshness rules decide which candidates may actually be supplied.
Only after that does generation begin. The prompt should identify the selected evidence and make clear that retrieved text is data, not trusted instructions. The returned claim should cite a supplied source ID, and application code should reject citations to evidence that was never provided. This trace separates two claims. “The answer looks right” is weaker than “this exact evidence path supports the answer.”
Retrieval quality is not one number
Different stages need different measurements. Recall@k asks whether relevant evidence appeared in the candidate set. Ranking metrics care where it appeared. Grounding or support checks ask whether answer claims are justified by supplied evidence. Citation checks ask whether source identities are valid. Access and freshness checks are constraints: a fluent answer cannot compensate for leaking an unauthorized document or using an expired policy.
This means a system change should be evaluated at the stage it is supposed to improve. If a new embedding model raises recall but answer quality stays flat, the bottleneck may have moved to reranking or generation. If answer quality drops while retrieval metrics are unchanged, investigate the later stages rather than assuming the index became worse.
How to study this level
Use small corpora first. For each query, write down the relevant source ID before running retrieval, then inspect candidate IDs and scores rather than looking only at the final answer. Change one variable at a time—chunk size, representation, hybrid weight, reranker, prompt—so the effect remains interpretable.
The browser labs make these transformations deterministic and inspectable. At L9.14, the optional real-model exercise replaces the toy representation with an actual sentence-transformer and can also call a real generator. The same debugging rule still applies: preserve the retrieval trace so a model call does not hide where the evidence came from.
Working rule
A correct final sentence does not prove the retrieval system worked. Keep retrieval evidence visible:
- which query was used;
- which chunks were eligible;
- which scores were produced;
- which chunks were supplied to generation;
- which source IDs support the final claims.
That separation is the main debugging tool for this level.
Completion is stored locally on this device.