본문으로 건너뛰기
L3.11

Pretrained Models

Goal

By the end of this lesson, you can describe a pretrained model as a reproducible contract—architecture, weights, preprocessing, label mapping, provenance, and evaluation mode—not merely a file of parameters.

Weights only make sense inside the computation that uses them​

Imagine a saved encoder whose first layer expects 64 input values and produces 32 hidden values.

If you build a new encoder with 24 hidden units and try to load the old weight tensor, the shapes no longer match.

That visible error is useful. It prevents one set of learned numbers from being silently interpreted under a different architecture.

But matching shapes are only the beginning.

The same weights can fail under wrong preprocessing​

Suppose source training normalized digit pixels into a particular range.

At reuse time, you add 2.0 to every normalized feature before passing it to the encoder.

The architecture and state dictionary still match. The model loads successfully. Yet the input distribution reaching the weights is now different from the one the encoder was trained to interpret.

This can shift activations and damage validation behavior.

So a pretrained-model contract includes at least:

  • architecture/configuration;
  • exact weight revision/checksum;
  • preprocessing and input shape;
  • label or vocabulary mapping when applicable;
  • training/source provenance when available;
  • evaluation mode and expected output semantics.

Inspect the notebook model contract​

The notebook trains a small source encoder on scikit-learn digits and reloads it.

  1. Open the notebook and run the code cell. It trains a small encoder on the digits data, saves its weights with saved = copy.deepcopy(model.encoder.state_dict()), and reloads them into a fresh Encoder(32) called good.
  2. Read mean_feature_shift_from_wrong_preprocessing=.... The notebook feeds the same validation images twice: once correctly (val) and once with every pixel shifted up by 2.0 (val+2.0). The weights never changed, yet the encoder's features moved by more than 0.1 on average. Wrong preprocessing alone changed what the model “sees.”
  3. Read shape_mismatch_detected=True. The notebook deliberately tried to load the saved 32-unit weights into Encoder(24). PyTorch refused with a shape error, and the notebook recorded that refusal instead of hiding it.
  4. Now make the preprocessing mistake smaller. Find good(val+2.0) and change 2.0 to 0.5. Before running, predict whether the feature shift should grow or shrink. Run the cell and compare.
  5. Change the offset back to 2.0. The lesson is that the checkpoint, its input preprocessing, and its architecture form one contract: change any one of them and the saved weights no longer mean what they did.

Loading lab…

Loading successfully is necessary, not sufficient​

A model can load cleanly and still be used incorrectly because:

  • class IDs changed order;
  • tokenizer/vocabulary mapping changed;
  • pixel normalization changed;
  • evaluation mode differs;
  • an incompatible data population is being treated as equivalent.

Successful deserialization only proves that certain structural checks passed.

Provenance is part of engineering evidence​

For an externally released model, reproducibility and responsible use may also require:

  • repository revision;
  • license;
  • training-data description where available;
  • known limitations;
  • intended/unsupported uses.

A pretrained model is an artifact with a history and input/output assumptions.

A pretrained weight file is not a complete model description​

Suppose you download an encoder checkpoint.

To reproduce its behavior, you may also need:

  • the exact architecture;
  • input size and channel order;
  • normalization constants;
  • label mapping for the original task;
  • library/version assumptions.

A tensor file can load successfully while the surrounding preprocessing is wrong.

Provenance answers “what exactly am I reusing?”​

Useful provenance records include where the checkpoint came from, which version it is, what data/task it was trained for, and any license or usage constraints.

This matters technically as well as administratively. A similarly named checkpoint from another release may have different preprocessing or parameter shapes.

Treat pretrained artifacts like dependencies with identities, not anonymous bags of useful weights.

Quick Check

1. What is part of a pretrained-model contract?
2. What should an incompatible parameter shape do?
3. Can correct weights still produce bad results?

0 of 3 questions answered.

Key Takeaways

  • A pretrained model is more than a weight file.
  • Architecture and parameter shapes must match.
  • Preprocessing and label/vocabulary mapping are part of model meaning.
  • Provenance, revision, and limitations support reproducible use.
  • Successful loading does not prove correct evaluation.

Next Lesson

Next, you will allow the pretrained representation itself to move and measure whether fine-tuning helps or damages the target task.

References

Lesson actions

Completion is stored locally on this device.

View progress