Skip to main content
L15.2

Golden Sets and Regression Suites

Goal

Build a reviewed set of representative and high-value cases that can catch regressions without pretending the set is a complete model of the real world.

Turn past behavior into reviewed cases​

A regression is a failure that returns after a change. The change might be a new model, a prompt edit, a retrieval update, a tool schema change, or a different serving configuration. Without repeatable cases, teams often discover regressions through users rather than tests.

A golden set is a curated collection of cases with reviewed expectations. “Golden” does not mean perfect or eternal. It means the cases are important enough that their inputs, provenance, expected behavior, and review history are controlled. When the product changes, the set gives you a stable comparison point.

Consider a support assistant. A weak golden set might contain fifty easy FAQ questions copied from one help page. A stronger set might include common questions, ambiguous wording, outdated-product references, permission-sensitive requests, retrieval failures, refusal cases, and examples that previously caused incidents. The stronger set is smaller than the real world, but it better represents decisions the product must get right.

Give every case identity and provenance​

Each case needs an identity. Store the input, relevant context, expected behavior, and tags such as refund, account-access, or citation-required. If the expected answer is not exact text, store an observable rule instead: “must mention the 30-day limit,” “must not claim the refund already happened,” or “must ask for approval before the side effect.”

{"case_id":"order-status-042","source":"incident-2026-08","expected":"state uncertainty clearly","tags":["orders","regression"]}

Provenance matters because test data can become misleading. Record where a case came from and why it belongs in the suite. A case copied from production may need privacy review or de-identification. A synthetic case should be labeled so reviewers do not mistake it for observed traffic. A case based on an incident should link to the failure category, not to secret user data.

Cover important slices and fix bad tests​

Regression suites also need coverage across important slices. If a system supports five task families but the set contains 90% examples from one family, the aggregate pass rate mostly describes that family. A simple coverage table can reveal imbalance before any model is run.

Do not freeze incorrect expectations just because a case is old. Golden sets require maintenance. Product policy can change, source documents can be corrected, and a previously accepted answer may turn out to be wrong. Updating an expectation should be a reviewed change with a reason, because changing the test at the same time as the model can hide a regression.

A useful workflow separates candidate failures from test defects. When a new version fails a case, first inspect the evidence. If the candidate is wrong, fix the system and keep the case. If the expected behavior is outdated, fix the case with review history. If the case is ambiguous, rewrite or split it rather than forcing one arbitrary label.

Add fresh cases and compare case-level regressions​

You also need fresh cases. A suite made entirely from known failures can overfit the development process. Hold out some cases, rotate newly sampled cases, or periodically add reviewed examples from new product behavior. The purpose is not to surprise the team; it is to prevent the test collection from becoming a narrow script that every release learns to pass.

The regression record should report both overall results and case-level changes from the baseline. “97% pass” is less informative than “two newly failing account-access cases and one fixed citation case.” The second description points directly to what changed.

HELM's broader lesson applies again: evaluation coverage should be explicit. A golden set is one scenario-specific evaluation asset, not proof that a model is generally safe or capable. You can say what the suite covers, how it was built, and which failures it would catch. You should also say what it does not cover.

Predict

A team changes both the model and three expected answers on the same day, then reports no regression. What is the main review problem?

Run the local Lab​

Run:

python3 labs/notebooks/level-15/l15-02-golden-regression.py

The Lab compares a baseline and candidate by case ID and slice.

  1. Run it unchanged. Record the baseline/candidate pass rates, newly_failing, fixed, and failures by slice.
  2. In CANDIDATE, find {"id":"auth-1","slice":"authorization","passed":True}. Before editing, predict which case should appear in newly_failing and which slice should gain one failure if only this result flips.
  3. Change only that candidate value to "passed":False, then rerun.
  4. Confirm auth-1 appears in newly_failing, the authorization slice records one failure, and the already-fixed grounding case remains fixed. Explain why the case-level diff is more informative than one pass-rate number.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-15/l15-02-golden-regression-exercise.py

Implement case-ID matching, newly-failing/fixed case detection, and failure counts by evaluation slice. The exercise intentionally includes both a regression and a fix.

Run:

python3 labs/notebooks/level-15/l15-02-golden-regression-exercise.py

The starter intentionally stops at TODO until you implement the missing logic. A correct solution reaches the final PASS: marker. Use the solved deterministic Lab as a comparison only after your own attempt.

Quick Check

1. Which description best fits a golden set?
2. Why tag cases by task or risk slice?
3. What should happen when an expected result is discovered to be wrong?

0 of 3 questions answered.

Explain it back​

Design five cases for one AI feature. Give each a case ID, a reason it belongs, one useful slice tag, and an observable expected behavior that does not depend on exact wording.

Key Takeaways

  • Golden sets are reviewed, controlled, and maintained.
  • Preserve provenance and observable expectations for each case.
  • Report case-level and slice-level regressions, not only one pass rate.
  • Review expectation changes independently from candidate changes.
  • Add fresh evidence so the suite does not become a narrow development script.

Next Lesson

Next, L15.3 — Human Evaluation covers qualities that deterministic checks cannot fully judge.

References

Lesson actions

Completion is stored locally on this device.

View progress