Skip to main content
L8.6

LoRA Intuition

Goal

Explain LoRA as a low-rank learned update to a frozen weight matrix, trace the update shapes, and compare parameter counts for different ranks.

Start with the problem before the matrices. Suppose a large layer already performs useful work, but you want training to learn a small correction without rewriting every value in that layer. LoRA keeps the original weight matrix fixed and learns a compact correction path beside it.

A useful picture is a large sheet that would be expensive to edit directly. Instead of storing another full-size sheet of corrections, LoRA describes the correction using two thinner tables that multiply together to produce a full-size update when needed. The technical word low-rank describes that structural shortcut. The matrices below are the precise way to express the shortcut, not a different idea from it.

Remember from L2.4 — Layers and Shapes: matrix shapes tell you which dimensions can connect. In LoRA, the small rank dimension r is the narrow middle dimension that lets two thinner matrices rebuild an update with the same outer shape as W.

Suppose a linear layer has a frozen weight matrix:

W shape = (out_features, in_features)

A normal full fine-tune learns a new value for every element of W. LoRA keeps W frozen and learns a smaller update built from two matrices.

Factor the update​

A common notation is:

ΔW = B @ A

with:

A shape = (r, in_features)
B shape = (out_features, r)

Then:

B @ A shape
= (out_features, r) @ (r, in_features)
= (out_features, in_features)

So ΔW has exactly the same shape as W. The adapted layer behaves conceptually like:

y = W x + scale * (B A) x

The base contribution remains. The low-rank branch adds a learned correction.

Rank controls adapter capacity​

Take:

in_features = 8
out_features = 6
rank r = 2

Full matrix parameters:

6 × 8 = 48

LoRA parameters:

A: 2 × 8 = 16
B: 6 × 2 = 12
total = 28

For this tiny matrix, the saving is modest. But for a 4096 × 4096 matrix with rank 8:

full: 4096 × 4096 = 16,777,216
LoRA: 8×4096 + 4096×8 = 65,536

That is a dramatically smaller trainable update.

Low rank is a structural assumption​

Choosing rank r assumes the useful adaptation can be represented through a relatively low-dimensional update. A larger rank increases adapter capacity and parameter count. A smaller rank is cheaper but may be too restrictive for some adaptation. Rank is therefore not “quality level.” It is a capacity/resource hyperparameter that needs evaluation.

Scaling changes update magnitude​

Many LoRA implementations multiply the update by a factor related to alpha/r. The exact implementation should be checked in the library/model you use. Conceptually, scaling controls how strongly the learned low-rank branch contributes relative to the frozen base path. A rank change without checking scale can alter both capacity and effective update magnitude.

LoRA can target selected matrices​

You do not have to adapt every matrix. Common implementations target selected projection layers such as attention projections, depending on architecture and experiment design. This creates another reproducibility requirement:

rank
alpha/scale
target modules
dropout if any
base model identity

“Used LoRA” is not enough detail to reproduce an adapter.

Rank limits what update patterns can be represented directly​

Rank is easier to understand by looking at an extreme case. If rank is 1, then:

ΔW = B @ A

is built from one column in B and one row in A. Every row of the update is therefore a scaled version of the same row pattern from A. That is a strong structural restriction. With a larger rank, several such patterns can be combined.

This does not mean rank 1 is always weak or rank 64 is always better. A narrow task may need only a small update subspace, while another task may need more capacity. The useful experiment is to compare rank under the same data and evaluation, then record both quality and parameter cost.

The adapter receives gradients through the frozen computation​

Freezing W does not stop gradients from flowing through the layer to trainable adapter parameters. During training, the loss depends on the combined output:

Wx + scale·B(Ax)

The optimizer updates A and B while leaving W unchanged. So there are two separate questions:

  • does a parameter participate in the forward/backward computation?
  • does the optimizer update that parameter?

A frozen base participates in the computation even though its values remain fixed.

Predict

If W has shape (12, 8) and LoRA rank is 3, which shapes make B@A match W?

Run the Browser Lab​

The Browser Lab uses tiny list-based matrices.

  1. Click Run. Read A shape: (1, 2), B shape: (2, 1), and B@A: [[0.5, 0.5], [-1.0, -1.0]]. The product has the full 2 × 2 shape even though it came from two thin factors.
  2. Check one value of B@A by hand: row 2, column 1 is -2.0 × 0.5 = -1.0.
  3. Read the 8 × 8 lines: full matrix parameters: 64, then rank 1 → 16, rank 2 → 32, rank 4 → 64.
  4. Predict when LoRA stops saving parameters. Each rank adds 8 + 8 = 16 parameters.
  5. Change for rank in (1, 2, 4): to for rank in (1, 2, 4, 5): and click Run. Rank 5 needs 80 parameters—more than the full matrix. Low rank only saves parameters when the rank is small compared with the matrix dimensions.
  6. Press Reset afterward.

Loading lab…

Quick Check

1. What shape must ΔW have?
2. What happens when LoRA rank increases while layer dimensions stay fixed?
3. What does low rank assume?

0 of 3 questions answered.

Explain it back​

For a 100 × 80 weight matrix with rank 4, give the shapes of A and B, compute LoRA parameter count, compare with full-matrix parameter count, and explain what rank controls.

Key Takeaways

  • LoRA learns a low-rank update while the base matrix remains frozen.
  • A and B multiply to the same shape as W.
  • Rank controls adapter capacity and parameter count.
  • Scaling controls update contribution and must be recorded.
  • Target-module identity is part of reproducibility.

Next Lesson

Next, implement the low-rank branch directly and verify its shape, initialization behavior, and trainable parameter count.

References

Lesson actions

Completion is stored locally on this device.

View progress