Level 8 — Fine-Tuning and Model Adaptation
Level 7 treated a pretrained LLM as a fixed probabilistic component and improved reliability through prompts, context, validation, and evaluation. Level 8 asks a different question:
When is changing the model itself justified, and how do we do it without losing track of data, evaluation, cost, or regressions?
Learning goal
Build an adaptation workflow you can explain and audit:
task gap
→ adaptation decision
→ training examples
→ cleaning/formatting
→ supervised fine-tuning
→ parameter-efficient adaptation
→ LoRA
→ quantization trade-offs
→ preference data/objectives
→ before/after evaluation
→ forgetting/drift checks
→ model card + training record
What mastery looks like
By the end of the level, you should be able to:
- decide whether prompting, retrieval, or fine-tuning is the right lever for a specific failure;
- build instruction-tuning examples with clear input/output contracts;
- clean, deduplicate, split, and format training data without leaking evaluation information;
- explain supervised fine-tuning as next-token optimization on desired responses;
- distinguish full fine-tuning from parameter-efficient fine-tuning;
- trace the shapes and parameter count of a small LoRA update;
- explain what quantizing a frozen base model changes and what it does not;
- build and inspect preference pairs;
- explain preference optimization as increasing relative preference for chosen behavior;
- compare a base model and adapted model on the same fixed evaluation set;
- detect regressions, drift, and catastrophic forgetting on retained capabilities;
- package model identity, data identity, adapter/checkpoint identity, metrics, limitations, and intended use.
Mini checkpoint
Complete the checkpoint after 8.07. You should be able to trace the path from an adaptation decision through an instruction dataset and a low-rank adapter before moving to quantization and preference optimization.
Level project
After 8.14, complete Adapt a Small LLM. The default project path is intentionally small and deterministic. L8.14 includes an optional real-model LoRA exercise with PEFT; keep the same evaluation evidence and release checks when you run it.
The decision comes before the training run
Fine-tuning is tempting because it sounds like the most direct way to “teach” a model. In practice, changing weights is useful only for certain kinds of gaps. If the missing information changes every day, retrieval is usually a better fit. If a rule must never be violated, application code is stronger than a probabilistic weight update. If a prompt already produces the desired behavior reliably, training may add cost without solving a real problem.
This level therefore begins with diagnosis. You will name the behavior that is missing, define a fixed evaluation that exposes the gap, and compare fine-tuning with cheaper or more controllable interventions. Only after the intervention choice is justified do we move into datasets, supervised loss, parameter-efficient methods, and LoRA. That order prevents “we trained a model” from becoming the goal by itself.
Follow one artifact chain
Think of an adaptation run as a chain of identities. A base checkpoint is adapted using a particular dataset split and serialization template. The run has an optimization configuration and produces an adapter or checkpoint. That artifact is evaluated on target cases and on retention cases that represent behavior you do not want to lose. Finally, a model card and training record explain what was changed, what evidence was observed, and what remains untested.
If any identity in that chain is missing, a result becomes difficult to reproduce. An adapter without its compatible base model may be unusable. A strong metric without the exact evaluation set may be impossible to interpret. A “before versus after” comparison using different prompts can attribute an improvement to the wrong cause. The small deterministic labs in this level repeatedly expose these relationships so that lineage becomes a habit rather than paperwork added at the end.
A concrete mental model for LoRA
Suppose a layer contains a large frozen weight matrix. Full fine-tuning would update that whole matrix. LoRA instead learns a much smaller pair of matrices whose product forms an update with the same outer shape as the original weight. The base path still computes its normal result, and the low-rank branch adds a scaled correction. Rank controls the size of that trainable subspace; it does not mean the base weights disappear.
This mental model connects several later topics. PEFT is about limiting what is trainable. Quantization is about how values are represented in memory and computation. The two ideas can be combined, but they solve different problems. Evaluation then asks whether the small update changed the intended behavior without damaging capabilities that should remain stable.
How to study this level
Do not judge an adaptation from training loss alone. For every experiment, keep a small before/after table with the same held-out cases, the exact base identity, adapter configuration, and any retention metrics. When you change rank, data, or learning settings, change one factor at a time when possible.
The browser labs intentionally use lists and small arithmetic so the mechanics are inspectable. At L8.14, the optional real-model exercise attaches a genuine PEFT LoRA adapter to a small pretrained model. Use it to connect the toy math to real library objects. Keep the same discipline: a few successful training steps show that the mechanism works. They do not show that the adapted model is globally better.
Working rule
Fine-tuning is not a reward for having a model problem. It is one intervention among several. Before changing weights, define the failure. Record why prompting or retrieval is not enough. Name the training data that should change behavior and the evaluation cases that must not regress.
Completion is stored locally on this device.