Skip to main content
L9.14

RAG System Integration Workshop

Goal

Integrate corpus/version identity, chunking, eligibility, retrieval, reranking, grounding, citations, and stage-specific evaluation into one reviewable RAG workflow.

Follow one claim through the system: "XR-4172 has an 18-month warranty." First ask which source revision contains that fact. Then identify the chunk created from it, whether the current user is allowed to retrieve it, how it ranked for the query, whether it survived reranking, whether it was actually placed in the model context, and which citation the final claim used.

That chain is the workshop's main mental model. Every transformation should preserve enough identity to answer "where did this evidence come from, and was it still present here?" If a final answer is wrong, debug the earliest point where the correct evidence disappeared, became ineligible, was outranked, or was misused.

source documents
→ ingestion/versioning
→ chunks + metadata
→ eligible corpus
→ retrieval candidates
→ reranked context
→ grounded request
→ answer + claim citations
→ stage metrics

Define one traceable request​

For one query, record:

request_id
authenticated principal/tenant
original query
rewritten query if any
corpus/index version
eligible chunk count
candidate chunk IDs + scores
reranked chunk IDs + scores
supplied evidence IDs
generated structured answer
claim → source mappings
validation results

This trace lets another person reproduce the reasoning path without relying on a screenshot.

Release criteria should span stages​

A RAG change can improve one metric while harming another. Example release rule:

retrieval recall@5 >= 0.90
grounded support pass >= 0.90
citation identity pass == 1.00
missing-evidence abstention >= 0.95
ACL violations == 0
stale-source violations == 0

The exact thresholds depend on the application. The important point is that retrieval speed or answer fluency cannot compensate for an authorization violation.

Compare one change at a time​

Suppose hybrid search improves recall. Then you also change:

  • chunk size;
  • embedding model;
  • reranker;
  • prompt.

If the answer metric changes, the cause is ambiguous. Prefer staged experiments:

  1. retrieval method change with fixed chunks;
  2. reranker change with fixed candidate set;
  3. prompt change with fixed retrieved evidence.

System integration does not remove controlled experimentation.

Preserve source and index lineage​

A result should identify:

corpus snapshot
chunker version
embedding model
index build
lexical index
reranker version
prompt version
generator version
evaluator version

Without those identities, a later “same query” run may actually use a different evidence system.

Diagnose with the earliest failed invariant​

Example: Final answer cites the wrong warranty. Trace:

  1. current policy exists? yes;
  2. current policy chunk created? yes;
  3. user authorized? yes;
  4. chunk eligible? yes;
  5. chunk retrieved top-5? no.

Stop there. Do not tune the citation formatter. The earliest failure is retrieval. Another case:

1–5 pass; 6. correct chunk supplied? yes; 7. output claim says 24 months while source says 18.

Now the failure is grounded generation/support validation.

Deterministic local validation path​

The p09 project uses a deterministic local corpus and score fixtures. It validates that the learner can:

  • preserve stable chunk/source IDs;
  • apply tenant/access eligibility;
  • rank candidates;
  • validate supplied-source citations;
  • compute retrieval metrics;
  • reject stale/unauthorized evidence;
  • separate retrieval and answer failures.

A live embedding model/vector database/LLM can be added as Builder/Engineer evidence, but it does not replace the deterministic evidence requirements.

Predict

A final answer is wrong and the relevant chunk was never in top-k. Which fix should be investigated first?

Run the integration package​

  1. Run the p09 reference validator.
  2. Confirm the passing fixture succeeds.
  3. Run the intentional failure fixture containing an unauthorized/stale evidence path.
  4. Inspect the first invariant that fails.
  5. Complete your learner starter.
  6. Produce a run record with retrieval and answer metrics.
  7. If you connect a live model or vector database, preserve the reduced path and add its artifact identities.

Loading lab…

Optional real-model retrieval extension​

After the deterministic RAG package passes, replace the synthetic retrieval representation with a real sentence-embedding model:

pip install -r labs/real-model/requirements.txt
python labs/real-model/l09_embedding_rag_real_model.py

The script downloads a pretrained sentence-transformer, embeds a small corpus and query, ranks normalized vectors, and prints the retrieved source IDs and similarity scores. To also send those retrieved chunks to a real instruction-tuned generator, add --generate:

python labs/real-model/l09_embedding_rag_real_model.py --generate

Keep the retrieved IDs and scores beside the generated answer. The point is to preserve the same retrieval-versus-generation debugging boundary when real models replace the deterministic fixtures.

Quick Check

1. Why should RAG release criteria include both retrieval and grounded-answer metrics?
2. Which failure should have zero tolerance in the example release rule?
3. Which record is sufficient to reproduce a RAG run?

0 of 3 questions answered.

Explain it back​

Trace a single RAG request from source snapshot to final cited answer. Name one invariant at each boundary and identify which artifact versions must be recorded.

Key Takeaways

  • RAG is an evidence pipeline with multiple independent boundaries.
  • Release criteria should protect retrieval, grounding, citation, freshness, and access.
  • Change one system component at a time when diagnosing improvements.
  • Preserve corpus/index/model/evaluator lineage.
  • Debug the earliest failed invariant.
  • A deterministic local path can validate retrieval and grounding checks without requiring live infrastructure.

Next Lesson

Complete Evidence-Backed RAG. Level 10 will move from retrieval into tool use and function-calling workflows.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Evidence-Backed RAG

View progress