본문으로 건너뛰기
L8.4

PEFT and LoRA

Goal

By the end of this lesson, you can explain why PEFT trains a small component around a frozen base model and trace how LoRA builds a low-rank update with far fewer trainable parameters than a full matrix.

Imagine a large map that is already useful. Full fine-tuning redraws the map itself. A parameter-efficient method keeps the map and adds a smaller overlay that changes only what the new task needs.

That is the central idea of parameter-efficient fine-tuning (PEFT): most pretrained weights remain frozen while a much smaller parameter set is optimized.

Frozen means “not updated by the optimizer.” It does not mean removed, zeroed, or skipped. The base model still performs the forward computation.

Why PEFT can be cheaper​

Suppose a base has 1,000,000 parameters and an adapter adds 20,000 trainable parameters.

full fine-tuning trainable parameters: 1,000,000
PEFT trainable parameters: 20,000
loaded base + adapter: 1,020,000

The smaller trainable set can reduce gradient and optimizer-state cost. It can also make task-specific artifacts easier to store because several adapters may share one base checkpoint.

The adapter is not a complete model identity by itself. Reproducible use still needs the compatible base model and adapter configuration.

LoRA is one PEFT method​

LoRA keeps a base weight matrix W frozen and learns a correction delta W using two thinner matrices.

For a layer with shape:

W: (out_features, in_features)
A: (rank, in_features)
B: (out_features, rank)

the product B @ A has the same outer shape as W.

The adapted layer can be pictured as:

base output + scaled low-rank correction
W x + scale * B(Ax)

The original computation remains present. LoRA adds a learned correction path.

Rank controls adapter capacity and size​

For a 4096 by 4096 matrix with rank 8:

full matrix: 4096 * 4096 = 16,777,216 parameters
LoRA: 8*4096 + 4096*8 = 65,536 parameters

A larger rank gives the adapter more capacity but also more trainable parameters. A smaller rank is cheaper but may be too restrictive. Rank is a capacity/resource choice, not a universal quality score.

PEFT and full fine-tuning are different interventions​

Full fine-tuning directly updates the model parameters selected for training, often all of them. PEFT keeps most base values fixed and learns a smaller trainable component or subset.

That difference changes optimization memory, storage, artifact lineage, and sometimes deployment structure. It does not guarantee better behavior. PEFT can still overfit, learn bad labels, or cause regressions.

LoRA needs more than a rank number​

Record:

  • base model and revision;
  • rank;
  • alpha or scaling convention;
  • target module names;
  • dropout if used;
  • adapter identity;
  • evaluation results.

A configuration can even match zero intended modules if target names are wrong. Count trainable parameters and inspect matched modules before trusting a run.

Predict

If W has shape (12, 8) and LoRA rank is 3, which factor shapes make B@A match W?

Run the Browser Lab​

The Lab combines PEFT accounting with LoRA shapes.

  1. Complete trainable_fraction using adapter parameters divided by base-plus-adapter parameters.
  2. Inspect the LoRA factor shapes and confirm B @ A rebuilds the base matrix shape.
  3. Increase rank and predict how both trainable parameter count and trainable fraction change.
  4. Explain why the frozen base still matters even though it is not updated.

Loading lab…

Quick Check

1. What does frozen mean in PEFT?
2. What must the LoRA product B@A match?
3. What happens when LoRA rank increases while layer dimensions stay fixed?

0 of 3 questions answered.

Explain it back​

For a 100 by 80 weight matrix with rank 4, give A and B shapes, calculate LoRA parameter count, compare it with full-matrix fine-tuning, and explain which parameters remain frozen.

Key Takeaways

  • PEFT limits the parameter set that optimization updates.
  • A frozen base still performs the main forward computation.
  • LoRA represents a full-shape correction through two low-rank factors.
  • Rank trades adapter capacity against trainable parameter count.
  • Base identity, target modules, scaling, and evaluation belong to adapter provenance.

Next Lesson

Next, implement the low-rank branch directly and test its shape, initialization, scaling, and output invariants.

References

Lesson actions

Completion is stored locally on this device.

View progress