Skip to main content
L8.12

Catastrophic Forgetting and Drift

Goal

Recognize adaptation regressions on previously useful behavior, distinguish targeted improvement from broad drift, and design retention checks that reveal forgetting.

Fine-tuning pushes the model toward the new training objective. That can improve the target behavior while moving other behavior unintentionally. Two useful terms are: catastrophic forgetting — previously learned capability degrades substantially after new training; behavioral drift — outputs shift in a way that may be smaller, broader, or simply different from the intended adaptation.

The boundary between these labels is not always sharp. The important engineering job is to measure what changed.

A small target gain can hide a large old-task loss​

Imagine:

target-task accuracy:
base 70%
adapted 88%

general extraction:
base 91%
adapted 74%

The adaptation clearly improved its target. It also damaged a capability that may matter in the product. Calling the run “better” without the retention result would be incomplete.

Forgetting is about behavior, not parameter distance alone​

Two checkpoints can have a small parameter difference and a large output change on one sensitive slice. A large parameter difference can also leave a particular behavior mostly unchanged. Parameter movement is diagnostic evidence. Retention evaluation is behavioral evidence.

Do not substitute one for the other.

Data imbalance can pull behavior​

Suppose 95% of fine-tuning examples use terse one-line answers. Even if the stated goal is only domain terminology, the model may also learn a strong brevity bias. That is drift created by the dataset distribution. Inspect not only labels but also:

  • response length;
  • formatting;
  • tone/style;
  • domain mix;
  • missing-information frequency;
  • refusal/abstention rate.

Mitigation is an experiment, not a slogan​

Possible responses include:

  • reduce update magnitude/steps;
  • improve data balance;
  • mix retention examples into training;
  • lower adapter capacity;
  • change target modules;
  • early stop using held-out evidence;
  • reject the adaptation.

Which response works depends on the measured failure. Change one factor at a time where practical and rerun the same target/retention suites.

Track drift against the correct baseline​

If adapter v2 was trained from the original base, compare against that base and v1 separately when both matter. If v2 was trained on top of v1, the lineage is different. Artifact lineage affects interpretation. A training record should say exactly which model state each adaptation started from.

Drift becomes visible only when you define what should stay stable​

Adaptation is supposed to change behavior, so “the outputs changed” is not evidence of forgetting. Catastrophic forgetting or harmful drift means behavior that was previously useful has degraded beyond an acceptable boundary. That requires a retention set and a threshold chosen before looking only at the adapted model's best results.

Slice retention checks by capability when possible. A model might preserve extraction accuracy while becoming much worse at abstention, or keep task accuracy while drifting into an unwanted response style. If a mitigation is tested, keep the evaluation fixed so you can attribute the recovery to the mitigation rather than to a changed benchmark.

Predict

An adapter improves domain formatting but general extraction drops sharply. What is the strongest description?

Break and diagnose the Browser Lab model​

The Browser Lab contains a toy multi-task score table.

  1. Click Run once. regressions: [] is empty, but one check fails, because the TODO is not written yet.
  2. Read the scores. target_format rose from 0.70 to 0.88, but general_extraction fell from 0.91 to 0.74: the adapter got better at its target and forgot something else.
  3. Complete regression_flags: return every metric whose adapted score is lower than its base score by more than max_drop.
  4. Click Run again. You should see regressions: ['general_extraction']. abstention dropped only 0.01, which is inside the allowance.
  5. Change max_allowed_drop = 0.02 to max_allowed_drop = 0.0 and run. Now abstention is flagged too. The threshold is a policy choice, so write it down before you look at results.
  6. Restore 0.02. Now imagine a gentler adaptation (smaller updates or a more balanced training mix) and change general_extraction to {"base": 0.91, "adapted": 0.90}. Run again: no regressions are flagged.

Loading lab…

Quick Check

1. What is catastrophic forgetting evidence?
2. Why inspect response-style distributions in fine-tuning data?
3. Which evaluation should stay fixed when testing a forgetting mitigation?

0 of 3 questions answered.

Explain it back​

Give one example of a desired adaptation and one retained capability that could regress. Define a measurable threshold that would make you stop or revise the run.

Key Takeaways

  • Adaptation can improve the target while damaging older behavior.
  • Retention evaluation is stronger than parameter distance alone.
  • Data distributions can create unintended style or behavior drift.
  • Mitigations should be tested under fixed evaluation.
  • Model/adaptor lineage must be recorded.

Next Lesson

Next, package the evidence into a model card and training record that another person can audit.

References

Lesson actions

Completion is stored locally on this device.

View progress