Skip to main content

Level 9 project

Evidence-Backed RAG

Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.

Launch lesson: RAG System Integration Workshop

Prerequisite project: Adapt a Small LLM

Goal

Build a retrieval-augmented workflow whose evidence path is inspectable:

source snapshot
→ chunk/source identity
→ eligibility filter
→ retrieval ranking
→ supplied evidence
→ grounded output
→ citation/provenance validation
→ retrieval + answer evaluation
→ release decision

Task

  1. token_overlap
  2. eligible_source_ids
  3. retrieve
  4. validate_grounded_output
  5. validate_run

Validation

Run these commands from the downloaded Project folder or the public materials repository root.

python projects/tests/l09/validate_submission.py \
projects/starters/l09/rag.py \
path/to/your/rag-run.json

Rubric

Score each criterion from 0–4. A submission should score at least 16/20 overall and must not score 0 on access/provenance or evaluation.

Criterion4 — Strong evidence3 — Meets2 — Partial1 — Weak0 — Missing
Corpus/chunk/index lineageSource snapshot, chunker, embedding/retrieval identity, updates, and IDs are traceable and reproducibleClear lineage with minor gapsCore IDs exist but version/update evidence is incompleteMostly an unversioned document dumpNo trustworthy corpus/index identity
Retrieval and rerankingCandidate retrieval, exact/semantic trade-offs, cutoff, reranking, and failure slices are measured under fixed casesReproducible retrieval with minor gapsRetrieval works but diagnostics or slices are shallowOnly a few anecdotal queriesNo coherent retrieval system
Access and provenanceEligibility is enforced before ranking; supplied sources are authorized; supported answers cite supplied evidenceCorrect boundaries with minor gapsSome provenance checks but one important boundary is missingPrompt wording is the main security controlUnauthorized/unsupplied evidence can be accepted
RAG evaluation and failure analysisRetrieval metrics + grounded-answer metrics + abstention + earliest-boundary debug trace are all presentComplete evaluation with minor gapsFinal-answer evaluation present but retrieval diagnosis weakCherry-picked answer samples onlyNo fixed evaluation
Reproducibility and communicationCorpus/retriever/reranker/prompt/generator/evaluator identities, commands, limitations, and optional live-path differences are explicitReproducible package with small omissionsSeveral artifact versions or limitations missingMostly screenshots/output dumpsNo reproducible package

Objective checks

The reduced validator checks:

  • tenant/active eligibility filtering before ranking;
  • deterministic retrieval over eligible sources;
  • grounded-output provenance;
  • supplied-source subset relationships;
  • recall@2 and grounded-answer accuracy recomputation;
  • unauthorized-exposure count;
  • release-rule recomputation;
  • intentional cross-tenant failure rejection.

Human review

Human review should inspect whether:

  • chunking preserves the evidence needed for the task;
  • lexical/vector/hybrid choices match query types;
  • reranking improves the intended slices rather than only averages;
  • citations actually support claims;
  • freshness/deletion behavior is documented;
  • failures are diagnosed at the earliest broken boundary.

Assessment guardrail

A live vector database or model API is optional for repository acceptance. If provided, it extends the evidence but does not replace deterministic provenance, authorization, and evaluation checks.

← Back to all projects