Pretrained Models and Transfer Learning
Goal
By the end of this lesson, you can explain what transfer learning reuses, compare transfer fairly with training from scratch, and treat a pretrained model as a reproducible contract that includes architecture, weights, preprocessing, mappings, and where the model came from.
Reuse a learned representation instead of starting from zero
Suppose an encoder has already learned to recognize handwritten digits.
Now you have a related target task—classifying digits as loop-like or not—but only a small labeled target set.
Two useful starting strategies are:
- scratch: learn both the representation and target classifier from target examples;
- frozen transfer: reuse the pretrained encoder and train only a new target head.
Transfer learning asks whether features learned from the source task reduce what the target task must learn from limited data.
If source and target tasks depend on overlapping structure, reuse can help. If the source representation emphasizes the wrong distinctions, it can help little or even hurt. That is negative transfer.
A scratch baseline turns “pretrained worked” into evidence
A pretrained model producing a plausible score does not show that pretraining helped.
Keep fixed:
- the same target training examples;
- the same validation set;
- preprocessing;
- seed policy;
- metric.
Then compare frozen transfer with a scratch model under the same evidence.
The notebook uses scikit-learn's bundled digits dataset and a fixed seed of 13.
Part A — compare transfer with scratch
- Run the notebook from the top.
- Identify the source task: 10-class digit recognition.
- Identify the target task: loop-like digits
{0, 6, 8, 9}versus the other digits. - Record
frozen_transfer_accuracyandscratch_accuracy. - Find
train_size=40and change only that value to20. - Predict whether less target data could make the reused representation more valuable relative to scratch. Treat this as a hypothesis, not a promised result.
- Rerun from the setup cell and compare both accuracies on the same validation set.
- Restore
train_size=40.
Changing the target sample count and validation set together would make the comparison harder to interpret because more than one source of evidence changed.
Part A asks whether reuse helps. To make that result reproducible, another learner must also know exactly what was reused and how its inputs were prepared. That takes us from measuring transfer's effect to specifying the pretrained model's contract.
Pretrained weights are only part of the model
Now suppose the saved encoder expects 64 input values and produces 32 hidden values.
Loading those weights into an encoder with 24 hidden units should fail visibly because the parameter shapes do not match.
But matching shapes are not enough. The exact same weights can also behave incorrectly when the surrounding input contract changes.
A practical pretrained-model contract includes:
- architecture/configuration;
- exact weight revision or checksum;
- preprocessing and input shape;
- label, category, or vocabulary mapping when applicable;
- where the model came from and how it was trained, when known (provenance);
- evaluation mode and output semantics;
- license and known limitations for externally released artifacts.
A checkpoint is not an anonymous bag of useful numbers.
The same weights can see different inputs
The notebook reloads the trained encoder into a fresh Encoder(32) and passes the same validation images through two preprocessing paths:
- correct normalized input;
- the same input shifted by
+2.0.
The weights do not change, but the resulting features do.
Part B — inspect the pretrained contract
- Read
mean_feature_shift_from_wrong_preprocessing=.... - Read
shape_mismatch_detected=True. The notebook deliberately attempts to load the 32-unit weights intoEncoder(24), and PyTorch rejects the mismatch. - Change the wrong-preprocessing offset from
2.0to0.5. - Predict whether the measured feature shift should grow or shrink.
- Rerun and compare.
- Restore the offset to
2.0.
Loading lab…
The two parts belong together. Transfer learning is only reproducible when you know what was reused and the contract under which those reused parameters have meaning.
Loading successfully is necessary, not sufficient
A saved model can load without an error and still be used incorrectly because:
- class IDs changed order;
- vocabulary mapping changed;
- normalization changed;
- evaluation mode differs;
- an incompatible target population is being treated as equivalent.
Successful loading proves only that certain structural checks passed.
Do not let pretraining hide bad supervision
A pretrained encoder cannot repair shuffled target labels, data leakage, or an invalid evaluation split.
Pretraining changes the starting representation. Trust still comes from controlled target evidence.
A useful decision question is:
Does the source representation contain information the target needs, and does a fair target comparison show that reuse helps?
Quick Check
Key Takeaways
- Transfer learning reuses learned parameters and representations.
- Frozen transfer keeps the encoder fixed and learns a new target head.
- A scratch baseline tests whether reuse actually helps under the same target evidence.
- Pretrained weights require matching architecture, preprocessing, mappings, and a record of where the model came from.
- Successful loading does not prove correct model use.
- Pretraining does not fix bad labels, leakage, or unfair evaluation.
Next Lesson
Next, you will allow the pretrained representation itself to move and measure whether conservative fine-tuning helps or damages the target task.
References
- PyTorch, Transfer Learning for Computer Vision Tutorial.
- PyTorch, Saving and Loading Models.
Completion is stored locally on this device.