Skip to main content
L8.11

Evaluate Before and After Fine-Tuning

Goal

Compare base and adapted models on identical held-out cases, separate target-task gains from regressions, and report paired differences instead of cherry-picked generations.

Fine-tuning is a model change. A model change should be evaluated like an experiment:

same cases
same prompt/context policy
same decoding/evaluator
base model vs adapted model

If several other variables change at the same time, you cannot isolate the effect of adaptation.

Use paired cases​

Suppose five target-task cases produce:

case base adapted
A fail pass
B pass pass
C fail pass
D pass fail
E pass pass

Aggregate accuracy:

base = 3/5
adapted = 4/5

The average improved. But the table also shows a regression on case D. That regression identity matters.

Separate target and retention suites​

Use at least two suites:

Target suite

  • behavior the adaptation is supposed to improve.

Retention suite

  • behavior the base model should keep.

Example:

target:
domain-specific support formatting

retention:
general summarization
basic factual extraction
abstention behavior
structured-output compliance

An adaptation can improve target performance while harming retention. Without the second suite, the regression is invisible.

Keep stochastic evaluation controlled​

If generation is sampled, use repeated trials or a fixed seed policy. Record:

  • model/checkpoint/adapter identity;
  • prompt version;
  • case-set version;
  • decoding settings;
  • evaluator version;
  • retry policy.

A before/after comparison should differ primarily in the model adaptation being tested.

Use slices, not only one average​

Fine-tuning may improve common cases but harm long inputs, minority classes, rare formats, or abstention cases. Report slices tied to plausible failure mechanisms. For example:

short complete cases
missing-information cases
long cases
rare label
instruction-like source text

This turns “adapted model seems better” into a more precise statement.

Decide thresholds before looking at the result​

A release rule might be:

target success +8 percentage points or more
AND
retention regression <= 1 percentage point
AND
zero critical safety regression

The exact thresholds depend on the application. Writing them before looking at results reduces the temptation to move the goalposts after seeing a favored checkpoint.

Compute paired deltas, not only two independent averages​

For each case, define:

delta = adapted_score - base_score

For a binary pass/fail case, the most informative transitions are:

fail → pass improvement
pass → fail regression
pass → pass retained success
fail → fail unresolved failure

This paired view answers a different question from two separate averages. Two models could both score 80% while succeeding on different cases. An average-only report would call them tied even though 20% of cases may have changed in each direction.

Small evaluation sets have uncertainty​

If an adapted model improves from 8/10 to 9/10, that is one additional passing case. It may be a real improvement, but ten examples provide limited evidence about the broader task population. Do not hide small sample size behind percentages. Report counts and, for important decisions, expand the evaluation set or use an appropriate statistical uncertainty method.

For stochastic generation, repeated trials add another source of uncertainty. The evaluation protocol should state whether each case is run once, under a fixed seed, or across multiple samples.

Keep evaluator changes separate from model changes​

If you rewrite the rubric after seeing adapted outputs, the comparison can move even when model behavior does not. Version the evaluator or human-review instructions. A clean adaptation experiment changes the model while keeping the measurement boundary stable.

Predict

The adapted model improves the target suite but loses 15 points on a required retention suite. What should the report say?

Complete the evaluation Browser Lab​

The Browser Lab contains paired base/adapted case results.

  1. Click Run once. Every list in the report is empty, so the checks fail. The last line already shows the averages: pass rate base -> adapted: 0.6 -> 0.8.
  2. Complete compare_cases: a case is improved if it went from fail to pass, regressed if it went from pass to fail, and otherwise stays in unchanged_pass or unchanged_fail.
  3. Click Run again. You should see 'improved': ['A', 'C'] and 'regressed': ['D']. The average improved, but case D got worse.
  4. Add a retention regression: change case E to {"id": "E", "base": True, "adapted": False}.
  5. Click Run. The pass rates are now 0.6 -> 0.6—no average change at all—while regressed lists ['D', 'E']. The case-level report shows a trade that the average hides.
  6. Now apply a simple release rule: “no more than one regression.” Does this adaptation pass? Explain which regressed case you would investigate first and why. The Lab should report Result: experiment ran because the original paired-case baseline changed while the transition-classification invariants still pass.

Loading lab…

Quick Check

1. Why evaluate base and adapted models on the same cases?
2. What is a retention suite for?
3. Why report regressed case IDs beside averages?

0 of 3 questions answered.

Explain it back​

Design a target suite and a retention suite for one adaptation. Define one release rule before seeing the results, then explain why an average alone is insufficient.

Key Takeaways

  • Use paired before/after evaluation.
  • Keep target and retention suites separate.
  • Record all model/prompt/decoding/evaluator identities.
  • Inspect failure slices and regressed case IDs.
  • Predeclare acceptance thresholds when practical.

Next Lesson

Next, study catastrophic forgetting and behavioral drift as specific forms of adaptation regression.

References

Lesson actions

Completion is stored locally on this device.

View progress