Skip to main content
L8.14

Fine-Tuning Integration Workshop

Goal

Integrate adaptation choice, dataset identity, LoRA configuration, before/after evaluation, retention checks, and documentation into one reproducible package.

A fine-tuning experiment is complete only when another person can trace the whole chain:

observed gap
→ intervention decision
→ training data
→ split/format
→ adaptation config
→ output artifact
→ target evaluation
→ retention evaluation
→ release decision
→ limitations

The final workshop checks those relationships.

A measured gap comes first​

Example:

Base model succeeds on support content but fails the required diagnostic-label convention in 38% of held-out cases.

This statement gives fine-tuning a reason to exist. Without a measured gap, an adapted checkpoint is only “different.”

Tie every artifact to an identity​

A run package might contain:

base_model_id
dataset_id
split_id
template_version
adapter_config
adapter_id
target_eval_id
retention_eval_id
training_record
model_card

The validator should reject broken relationships. For example, if metrics claim adapter-v3 but the training record says adapter-v2, that is an evidence mismatch.

Keep the adaptation contract narrow​

A strong report might conclude:

On support-format-v3, adapter A improved exact-schema success from 78% to 91%. Retention-general-v2 changed from 95% to 94%, within the predeclared 2-point regression limit. The run does not establish quality on Korean input or contexts above 4k tokens.

That is stronger than:

Fine-tuning made the model much better.

The narrow claim states what improved and what it was compared with. It names the evaluation and the regression threshold. It also says what remains unknown.

Debug the first broken boundary​

Suppose the final adapted score is unexpectedly poor. Trace:

dataset rows
→ split membership
→ formatted examples
→ trainable parameter set
→ adapter config
→ checkpoint identity
→ evaluation prompt
→ metric output

Find the earliest mismatch. Examples:

  • validation examples leaked into training;
  • assistant-loss mask is wrong;
  • LoRA target module list is empty;
  • adapter is loaded onto the wrong base revision;
  • evaluation uses a different prompt template;
  • metric report points to another checkpoint.

Do not respond by changing rank, learning rate, and dataset simultaneously.

A deterministic local path can validate the workflow without expensive training​

The project includes a deterministic local validation path. It does not claim to reproduce the quality of a real LLM fine-tuning run. It does verify that the learner can:

  • decide why adaptation is justified;
  • preserve grouped train/eval boundaries;
  • compute LoRA parameter counts;
  • validate base/adapter compatibility metadata;
  • compare target and retention metrics;
  • reject a run with excessive retention regression;
  • package model-card/training-record evidence.

A real LoRA/QLoRA run can be added as Builder/Engineer evidence when compute is available.

Release decisions should be rule-based​

Example:

target improvement >= 8 points
retention regression <= 2 points
critical regression count == 0
lineage complete == true

A checkpoint that misses the rule may still be informative. It simply has not met the declared release condition. That distinction helps experiments remain honest.

Predict

An adapted checkpoint improves target cases but its metrics file names a different adapter than the training record. What failed?

Run the integration validator​

From the repository root, validate the passing reference package:

python projects/tests/l08/validate_submission.py projects/reference/l08/reference_solution.py projects/reference/l08/sample-run/adaptation-run.json

The first command should pass. Before running the failure fixture, compare only base_model_id with adaptation.compatible_base_id in that JSON and predict which invariant the validator should reject first. Then run:

python projects/tests/l08/validate_submission.py projects/reference/l08/reference_solution.py projects/reference/l08/failure-run/adaptation-run.json

Confirm it exits non-zero for the adapter/base lineage mismatch. After that, inspect the remaining recorded metrics as a second debugging exercise instead of changing several training settings at once.

After that, complete the learner project implementation and run the same validator on your own module and run record. If you perform a real model run, attach its identities and metrics without removing this deterministic reduced path.

Loading lab…

Optional real-model LoRA extension​

After the deterministic package passes, you can attach and train a real LoRA adapter on a small pretrained causal language model:

pip install -r labs/real-model/requirements.txt
python labs/real-model/l08_lora_real_model.py --steps 4

The script prints the base-model ID and trainable-parameter count. It masks prompt tokens from the supervised loss and trains only the adapter parameters for a few steps. It then compares one held-out response before and after, and saves the adapter under artifacts/real-model/. A GPU is recommended. Treat the run as a mechanics exercise. A few steps are not evidence that model quality improved. Use the target and retention evaluations from this level before making that claim.

Quick Check

1. What is the strongest reason to keep a reduced deterministic acceptance path?
2. What should a release rule contain?
3. Where should debugging begin when a packaged run is inconsistent?

0 of 3 questions answered.

Explain it back​

Trace one complete adaptation run from observed gap to release decision. Name every artifact identity needed to reproduce the claim and one failure that the validator should reject.

Key Takeaways

  • Fine-tuning is a reproducible experiment, not only a training command.
  • Data, base, adapter, evaluation, and documentation identities must agree.
  • Target gains and retention constraints belong in the same release decision.
  • Reduced deterministic checks validate engineering contracts without pretending to measure full model quality.
  • Debug from the earliest broken boundary.

Next Lesson

Complete Adapt a Small LLM. Level 9 will move from changing model weights to supplying external knowledge through retrieval-augmented generation.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Adapt a Small LLM

View progress