본문으로 건너뛰기
L6.8

Validation Loss and Overfitting

Goal

Read training and validation loss together, identify a simple overfitting pattern, and choose checkpoints using a stated held-out criterion.

Training loss answers: “How well do the current parameters fit the examples used for updates?” Validation loss asks a different question on text not used for updates.

Consider:

step 0 20 40 60 80
train 4.2 3.1 2.4 1.9 1.5
valid 4.3 3.3 2.8 2.9 3.2

The training curve keeps improving, but held-out loss is best near step 40. The final step is not automatically the best checkpoint.

Training keeps falling while validation turns upward

These are the same five checkpoints as the table above. Reveal them in order and watch when the two signals stop improving together.

Step 40 · train 2.400 · validation 2.800
Training loss Validation loss
watch heretraining steploss

Read the gap as evidence, not a verdict​

A widening train/validation gap is a warning pattern, not a magical overfitting detector. Validation can also worsen because the held-out set is tiny, the evaluation pipeline changed, or the training and validation text come from meaningfully different distributions.

That is why the correct sequence is:

  1. verify the held-out boundary and evaluation procedure;
  2. check whether the change is larger than ordinary measurement noise;
  3. compare multiple checkpoints under the same evaluation setup;
  4. only then interpret persistent train-improves/validation-worsens behavior as evidence of overfitting.

Checkpoint selection is also a decision rule. If you say “choose the lowest validation loss,” write that rule before seeing the final samples. Otherwise it is easy to favor whichever checkpoint produced the most appealing anecdote after the fact.

For small teaching runs, the exact best step can vary slightly across compatible environments. The important evidence is the relationship between curves and the documented selection criterion.

Training loss answers a different question from validation loss​

Training loss measures prediction error on examples used to update the model. Validation loss measures error on held-out examples that did not provide parameter updates.

It is normal for training loss to be lower because the optimizer directly adapts to the training set.

A pattern such as:

step train validation
100 3.2 3.4
500 2.4 2.6
1000 1.8 2.7

suggests that later training is improving fit to training data while held-out prediction has begun to worsen.

Do not choose the checkpoint from generated prose alone​

A later checkpoint may produce one impressive sample by chance while having worse held-out loss and more failure cases overall.

Likewise, the lowest-validation-loss checkpoint can still have unacceptable product-specific failures.

Checkpoint selection should combine a declared quantitative criterion with fixed behavioral evaluation cases rather than browsing samples until one looks good.

Predict

Training loss falls from 2.0 to 1.4 while validation loss rises from 2.3 to 2.9. What is the strongest first interpretation?

Verify the evaluation pipeline before changing the model​

  1. Click Run. The Lab picks the checkpoint with the lowest validation loss: best validation step: 40 with best validation loss: 2.8. After step 40, training loss keeps falling while validation loss rises to 2.9 and 3.2.
  2. Change only the last validation point: valid = [4.3, 3.3, 2.8, 2.9, 3.2] becomes valid = [4.3, 3.3, 2.8, 2.9, 2.6].
  3. Before running, predict whether the selected best step changes.
  4. Click Run. The best step becomes 80 with loss 2.6. The selection rule follows the validation evidence, so a single late measurement can change which checkpoint is kept. That is also why one noisy validation number deserves a second look before you trust it.
  5. Press Reset afterward.

Loading lab…

If validation suddenly worsens, first verify: correct held-out split, evaluation mode where relevant, same tokenizer/preprocessing, no gradient updates, and enough examples for a meaningful average. A bad evaluation pipeline can imitate a model problem.

Quick Check

1. Which signal should guide the stated best-checkpoint rule in this lesson?
2. What pattern is a common overfitting warning?
3. Why keep validation text outside updates?

0 of 3 questions answered.

Transfer the idea​

A run has flat train and validation loss. Give two plausible hypotheses other than overfitting, and state one controlled diagnostic for each.

Key Takeaways

  • Train and validation loss answer different questions.
  • Diverging curves can reveal overfitting.
  • Best-checkpoint selection needs a stated held-out rule.
  • Audit the validation pipeline before treating every curve change as a model failure.

Next Lesson

Next, L6.9 — Sampling Text turns a checkpoint's next-token logits into generated sequences.

References

Lesson actions

Completion is stored locally on this device.

View progress