Validation Loss and Overfitting
Goal
Read training and validation loss together, identify a simple overfitting pattern, and choose checkpoints using a stated held-out criterion.
Training loss answers: “How well do the current parameters fit the examples used for updates?” Validation loss asks a different question on text not used for updates.
Consider:
step 0 20 40 60 80
train 4.2 3.1 2.4 1.9 1.5
valid 4.3 3.3 2.8 2.9 3.2
The training curve keeps improving, but held-out loss is best near step 40. The final step is not automatically the best checkpoint.
Training keeps falling while validation turns upward
These are the same five checkpoints as the table above. Reveal them in order and watch when the two signals stop improving together.
Read the gap as evidence, not a verdict
A widening train/validation gap is a warning pattern, not a magical overfitting detector. Validation can also worsen because the held-out set is tiny, the evaluation pipeline changed, or the training and validation text come from meaningfully different distributions.
That is why the correct sequence is:
- verify the held-out boundary and evaluation procedure;
- check whether the change is larger than ordinary measurement noise;
- compare multiple checkpoints under the same evaluation setup;
- only then interpret persistent train-improves/validation-worsens behavior as evidence of overfitting.
Checkpoint selection is also a decision rule. If you say “choose the lowest validation loss,” write that rule before seeing the final samples. Otherwise it is easy to favor whichever checkpoint produced the most appealing anecdote after the fact.
For small teaching runs, the exact best step can vary slightly across compatible environments. The important evidence is the relationship between curves and the documented selection criterion.
Training loss answers a different question from validation loss
Training loss measures prediction error on examples used to update the model. Validation loss measures error on held-out examples that did not provide parameter updates.
It is normal for training loss to be lower because the optimizer directly adapts to the training set.
A pattern such as:
step train validation
100 3.2 3.4
500 2.4 2.6
1000 1.8 2.7
suggests that later training is improving fit to training data while held-out prediction has begun to worsen.
Do not choose the checkpoint from generated prose alone
A later checkpoint may produce one impressive sample by chance while having worse held-out loss and more failure cases overall.
Likewise, the lowest-validation-loss checkpoint can still have unacceptable product-specific failures.
Checkpoint selection should combine a declared quantitative criterion with fixed behavioral evaluation cases rather than browsing samples until one looks good.
Predict
Verify the evaluation pipeline before changing the model
- Click Run. The Lab picks the checkpoint with the lowest validation loss:
best validation step: 40withbest validation loss: 2.8. After step 40, training loss keeps falling while validation loss rises to2.9and3.2. - Change only the last validation point:
valid = [4.3, 3.3, 2.8, 2.9, 3.2]becomesvalid = [4.3, 3.3, 2.8, 2.9, 2.6]. - Before running, predict whether the selected best step changes.
- Click Run. The best step becomes
80with loss2.6. The selection rule follows the validation evidence, so a single late measurement can change which checkpoint is kept. That is also why one noisy validation number deserves a second look before you trust it. - Press Reset afterward.
Loading lab…
If validation suddenly worsens, first verify: correct held-out split, evaluation mode where relevant, same tokenizer/preprocessing, no gradient updates, and enough examples for a meaningful average. A bad evaluation pipeline can imitate a model problem.
Quick Check
Transfer the idea
A run has flat train and validation loss. Give two plausible hypotheses other than overfitting, and state one controlled diagnostic for each.
Key Takeaways
- Train and validation loss answer different questions.
- Diverging curves can reveal overfitting.
- Best-checkpoint selection needs a stated held-out rule.
- Audit the validation pipeline before treating every curve change as a model failure.
Next Lesson
Next, L6.9 — Sampling Text turns a checkpoint's next-token logits into generated sequences.
References
- Radford et al., Language Models are Unsupervised Multitask Learners.
Completion is stored locally on this device.