본문으로 건너뛰기
L3.14

Adaptation Review: Choose, Fine-Tune, Explain

Goal

By the end of this lesson, you can choose among scratch training, frozen transfer, and fine-tuning using a controlled scorecard that includes quality, trainable parameters, representation movement, failure slices, and reproducibility.

Adaptation is an experimental choice, not a prestige ranking​

Three strategies can solve the same target task:

  • scratch: learn encoder and head from target data;
  • frozen transfer: reuse encoder and learn only a head;
  • fine-tune: start from transferred weights and update encoder too.

The most advanced-sounding option is not automatically best.

Your decision should depend on target evidence and engineering tradeoffs.

Make the strategy the main changed factor​

A fair comparison keeps fixed:

  • target train/validation examples;
  • preprocessing;
  • seed policy;
  • metric;
  • evaluation set.

Then compare:

EvidenceScratchFrozen transferFine-tune
validation quality???
trainable parametersmanyfewmany/some
encoder movementlearned from zerozeromeasured
failure slices???
compute/reproducibility cost???

A scorecard makes it harder to choose based on one attractive accuracy number.

Run all three strategies under the same target evidence​

  1. Open the notebook. Before running, confirm that all three strategies train on the same small target set: a and b come from small, which holds train_size=40 examples, and all three are scored on the same val.
  2. Run the cell. It prints one table: strategy | accuracy | trainable_parameters | encoder_movement.
  3. Compare the trainable_parameters column. The frozen model trains only its small head, while scratch and fine-tuning train the whole network.
  4. Read encoder_movement for fine_tuned. It shows how far fine-tuning moved the pretrained encoder; the other two rows show 0.000 because that measurement does not apply to them.
  5. Now change only the amount of target data. Find train_size=40 and change it to train_size=12.
  6. Before rerunning, predict which strategy should suffer most with fewer target labels. (Hint: which one has no pretrained knowledge to fall back on?)
  7. Run the cell again and compare the whole table with your first run. The notebook's check requires every strategy to stay above 0.75; with only 12 labels a strategy may fall below that and raise an AssertionError. That is evidence, not a bug in your edit—record which strategy failed.
  8. Restore train_size=40 afterward.

Loading lab…

Similar quality can justify the simpler adaptation​

Suppose:

  • frozen transfer: 91.5% validation accuracy;
  • fine-tuning: 91.8%;

If fine-tuning trains far more parameters, takes longer, moves the encoder substantially, and provides no meaningful robustness improvement, frozen transfer may be the better engineering choice.

The 0.3-point difference is evidence, but it is not the only relevant evidence.

Different evaluation conditions destroy the comparison​

If scratch uses one split and fine-tuning uses an easier split, the resulting table may look precise but cannot isolate strategy effect.

Audit the experiment before tuning the models.

State a decision with limits​

A useful conclusion might say:

“Under this 40-example target split, frozen transfer is preferred because it matches fine-tuning within the observed validation variation while training fewer parameters and keeping the encoder fixed. We would retest the choice if target data grows or the failure-slice gap changes.”

That is stronger than “frozen is best.” It says what evidence supports the decision and what could change it.

Make the choice from a scorecard, not from the method name​

Suppose you compare three strategies:

scratch
frozen transfer
fine-tuning

A useful scorecard can include validation quality, trainable parameter count, training time, representation movement, and important failure slices.

You may find that fine-tuning improves average accuracy slightly but costs far more compute and worsens one important slice. Or frozen transfer may nearly match it with much lower operational cost.

The lesson is not that one strategy should always win. The lesson is to make the decision traceable to evidence.

State the conclusion narrowly​

Prefer:

On this target dataset and validation split, fine-tuning improved the chosen metric by X while changing Y parameters.

Avoid:

Fine-tuning is better.

A narrow conclusion records the conditions under which the evidence was collected and makes it easier to revise the decision when the data, model, or operational constraints change.

Quick Check

1. What makes an adaptation comparison fair?
2. What is a reason to prefer frozen transfer?
3. What should accompany an overall accuracy?

0 of 3 questions answered.

Key Takeaways

  • Scratch, frozen transfer, and fine-tuning are strategies to compare, not a universal ranking.
  • Hold target evidence and evaluation conditions fixed.
  • Compare quality with trainable parameters, encoder movement, robustness, and reproducibility cost.
  • A simpler adaptation can be preferable when quality is effectively similar.
  • State decisions narrowly enough that new evidence can revise them.

Next Lesson

Complete Transfer Learning Across a Real Dataset. Then Level 4 turns raw text into model-ready tokens and embeddings.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Transfer Learning Across a Real Dataset

View progress