Mini Checkpoint — SFT, PEFT, and LoRA
Complete this checkpoint after L8.7.
1. Choose the intervention
For each problem, choose prompting, retrieval/tooling, application logic, or fine-tuning and justify the choice:
- a policy changed this morning;
- JSON must contain exactly three required fields;
- a model repeatedly uses the wrong domain-specific response convention despite a clear prompt;
- refunds above a fixed amount require human approval.
2. Design instruction data
Create one easy example, one decision-boundary example, and one abstention example for a task of your choice. State what distinct behavior each example teaches.
3. Prevent leakage
You have five rows from two conversations:
case-A turn-1
case-A turn-2
case-A turn-3
case-B turn-1
case-B turn-2
Explain why a row-random split can leak information and give a group-safe split.
4. Mark SFT loss positions
For a formatted sequence containing system instruction, user input, assistant marker, response, and EOS, mark which positions contribute under response-only loss. Explain why masked prompt tokens still affect the response.
5. Count PEFT parameters
A base layer has shape 512 × 512. LoRA rank is 8. Calculate:
- full matrix parameter count;
- A shape;
- B shape;
- total LoRA parameter count;
- LoRA/full ratio.
6. Trace a LoRA update
Explain:
x → Ax → B(Ax) → scale → add to Wx
Then state what should happen when:
- adapter scale is zero;
- the initial low-rank product is zero.
Check your reasoning after you try
- A policy changed this morning → prefer retrieval/tooling. Exact JSON shape → application logic or constrained validation. A stable response convention that prompting does not fix → fine-tuning may fit. A refund approval threshold → application policy, not model weights.
- Split related conversation turns as a group. Putting
case-Aturns in both train and validation leaks near-duplicate context. - Under response-only SFT loss, response tokens (and usually the terminating EOS) contribute to loss. Prompt tokens can be masked from loss while still influencing response hidden states.
- For a
512 × 512layer, the full matrix has 262,144 parameters. With rank 8, LoRA usesA: 8 × 512andB: 512 × 8, for 8,192 trainable parameters, or 3.125% of the full matrix. - If adapter scale is zero, the LoRA branch contributes nothing. If the initial low-rank product is zero, the base output starts unchanged while training can still update the adapter factors.
Pass condition
You are ready for L8.8 when you can justify the adaptation decision, protect the data/evaluation boundary, explain SFT loss masking, and trace a LoRA update from shapes to output.
Completion is stored locally on this device.