Skip to main content
L8.8

Quantization for Fine-Tuning

Goal

Explain quantization as a lower-precision representation of model weights, distinguish quantized frozen base weights from trainable adapters, and measure approximation error on a small example.

Imagine recording temperatures when your measuring device can store only a small set of values. A reading of 19.8°C might be stored as 20°C. The stored value is close and cheaper to represent, but it is an approximation rather than the original measurement.

Quantization applies this general idea to model numbers. A model can keep the same number of weight positions while representing those values with fewer bits or a restricted numerical format. That distinction matters: fewer bits is not the same thing as fewer parameters. In quantized adapter training, the compact base can stay frozen while separate adapter parameters remain trainable.

Quantization approximates values​

Suppose the original weights are:

[-1.0, -0.2, 0.3, 0.95]

A very small toy quantizer might allow only:

[-1.0, -0.5, 0.0, 0.5, 1.0]

Then the values become approximately:

[-1.0, 0.0, 0.5, 1.0]

The representation uses a smaller set of possible values, but introduces error. For each value:

error = quantized - original

Compression is therefore a trade-off, not a free transformation.

Lower precision is not the same as fewer parameters​

A model can still have the same number of weight positions. Each position is represented with fewer bits or a quantization scheme. Do not confuse:

  • parameter count — how many learned scalar positions exist;
  • precision/storage — how many bits or what representation stores them.

LoRA reduces the number of trainable adaptation parameters. Quantization reduces memory/storage/computation cost of represented weights.

QLoRA-style training separates roles​

A common idea in quantized adapter fine-tuning is:

quantized base weights → frozen
LoRA adapter weights → trainable at suitable precision

The adapter learns updates while the compressed base provides the pretrained computation. The exact formats and kernels are implementation details that vary across libraries/hardware. The durable concept is the separation between a memory-efficient frozen base and a trainable adapter.

Quantization can change outputs​

If weights are approximated, logits can move. A small weight error may be harmless in one case and change a boundary decision in another. That is why evaluation should include:

  • before/after task metrics;
  • sensitive boundary cases;
  • numerical sanity checks;
  • memory/resource measurements.

Do not infer behavior quality only from “the model loaded successfully.”

Memory claims need a denominator and scope​

If you report a memory saving, specify what is included:

  • base weights only?
  • optimizer state?
  • adapter weights?
  • activations?
  • gradients?
  • runtime overhead?

A four-bit weight representation does not mean the entire training process uses exactly one quarter the memory of a 16-bit full fine-tune. System memory includes many components.

Quantization changes representation, not the number of model decisions​

Quantization approximates values with a smaller or lower-precision representation. That can reduce the memory used by frozen weights, but it introduces approximation error and does not automatically shrink every other part of training memory. Activations, adapter parameters, gradients, optimizer state, and temporary buffers still matter.

This is why “4-bit means exactly four times less training memory” is too simple. Measure the system you actually run. In QLoRA-style training, a quantized base can remain frozen while higher-precision adapter parameters are optimized. The learner should be able to point to which values are quantized, which parameters are trainable, and which evaluation detects quality loss caused by approximation.

Predict

What does quantizing a frozen base model primarily change?

Run the Browser Lab​

The Browser Lab implements a tiny nearest-level quantizer.

  1. Click Run once. The quantized values are [-1.0, 0.0, 0.5, 1.0], but mean absolute error: 0.0 is wrong, so one check fails.
  2. Complete mean_absolute_error: average the absolute difference between each original value and its quantized value.
  3. Click Run again. You should see mean absolute error: 0.1125, and every check should pass.
  4. Make the levels finer: change coarse_levels to [-1.0, -0.75, -0.5, -0.25, 0.0, 0.25, 0.5, 0.75, 1.0].
  5. Before running, predict whether the error should shrink.
  6. Click Run. The error drops to 0.0375. The Lab should report Result: experiment ran: the coarse-level baseline changed, but the invariant checks still confirm nearest-level quantization, correct error computation, and unchanged parameter count.
  7. Notice that the number of weights never changed—still four. Finer levels need more bits to store each weight: five levels fit in 3 bits, nine levels need 4. Quantization trades storage per weight against approximation error, not parameter count.
  8. Press Reset afterward, after copying your function if you want to keep it.

Loading lab…

Quick Check

1. What is the difference between quantization and PEFT?
2. In a QLoRA-style setup, which component is typically trainable?
3. Why is '4-bit means exactly 4× less training memory' too strong?

0 of 3 questions answered.

Explain it back​

Explain the difference among parameter count, trainable parameter count, and weight precision. Then describe why quantization and LoRA can be used together without meaning the same thing.

Key Takeaways

  • Quantization approximates weights with lower-precision representations.
  • Approximation introduces numerical error that should be measured.
  • Parameter count and numerical precision are different quantities.
  • Quantized base weights can remain frozen while adapters are trained.
  • Resource claims should state what memory components are included.

Next Lesson

Next, move from demonstration targets to pairwise preference evidence: chosen and rejected responses to the same context.

References

Lesson actions

Completion is stored locally on this device.

View progress