Diffusion Models: Learning to Denoise
Goal
By the end of this lesson, you can explain diffusion generation as learned iterative denoising and compare its generation loop with autoregressive next-token generation.
Level 6 has followed one modern generative pattern so far:
Autoregressive generation
token → next token → next token → ...
Now consider a different job.
Suppose you have a clean low-resolution pattern:
█ · █ ·
· █ · █
█ · █ ·
· █ · █
Add a little random corruption and the structure is still visible. Add more and more corruption and eventually the original pattern is hard to recognize.
A diffusion model learns the opposite skill: given a noisy state, predict how to remove some of the noise.
Generation then runs that learned repair process repeatedly:
Diffusion generation
noise → less noise → less noise → ... → sample
The important idea comes before any architecture details.
Step 1: start with clean data
Training begins with ordinary examples from the data distribution: images, audio representations, or other continuous data.
Call one clean example x₀.
At this point there is no generation yet. We already have the clean training example.
Step 2: add noise in controlled steps
Imagine corruption levels from 0 to 4:
step 0 clean
step 1 slightly noisy
step 2 noisier
step 3 heavily noisy
step 4 almost entirely noise
The forward noising process is deliberately easy to run. We control how much noise is added at each step.
This process is not the generative direction. It creates training situations where the model can practice recognizing what corruption was added.
Predict
Step 3: train a model to predict what should be removed
For a noisy example, the model receives information such as:
- the corrupted sample;
- how far along the noising process it is;
- optionally, a conditioning signal.
A common training target is the noise that was added.
clean example
↓ add known noise
noisy example
↓ denoiser
predicted noise
↓ compare
actual added noise
If the predicted noise is wrong, the loss is larger. Optimizer updates gradually improve the denoiser.
Only now, after seeing the corruption-and-repair job, do we need a small amount of math. A toy corruption rule can be written as:
corrupted = (1 - noise_level) × clean + noise_level × noise
This linear blend is only a teaching rule for making the corruption visible. It is not the exact forward equation used by DDPMs. Real diffusion systems usually draw random noise from a defined distribution and combine it with the clean example using timestep-dependent coefficients.
A simple training loss can compare predicted noise with the known noise:
loss = mean((predicted_noise - added_noise)²)
Real diffusion systems use more carefully designed noise schedules and parameterizations. You do not need scheduler notation, stochastic differential equations, or a particular neural-network architecture to understand the core training signal.
Step 4: generate by starting from noise
At generation time there is no clean answer to reveal.
Instead:
- start from random noise;
- ask the trained denoiser what noise-like component should be removed;
- move to a slightly cleaner state;
- repeat for several steps.
That is why sampling is iterative.
One denoising step usually does not produce a final image-like sample any more than one next-token decision produces a whole paragraph.
Run a visible corruption-and-repair experiment
The Lab uses a four-number pattern instead of an image so every value remains inspectable.
It contains:
- one clean pattern;
- one fixed noise vector, standing in for one reproducible random draw;
- five forward corruption levels;
- a precomputed sequence of idealized reverse corrections;
- a learner-controlled
denoise_strength.
The precomputed corrections are deliberately perfect teaching fixtures: each one points from the current toy state toward a cleaner state. They are written directly into the fixture rather than calculated from the clean target during reverse generation. The clean pattern is used afterward only to measure error. These corrections are not an implementation of a DDPM's learned noise-prediction network. This Lab isolates the reverse-process intuition and does not claim to train a production image model in the browser.
Loading lab…
Use this sequence:
- Click Run unchanged.
- Read the forward errors from clean. They should grow as the noise level increases.
- Read the reverse errors from clean. They should shrink as denoising proceeds.
- Confirm that the reverse process starts from the fixed noise vector, which stands in for one reproducible random start, not from the clean target.
- Change only
denoise_strength = 1.0to0.5. - Predict what will happen: each correction is smaller, so after the same number of steps the final state should remain farther from the clean pattern.
- Run and compare the final error.
- Reset.
This is a process experiment. Notice the important boundary: run_reverse starts from the fixed noise vector and consumes only the prewritten corrections plus denoise_strength; it does not consult clean. A real diffusion model would produce those updates with a trained denoiser rather than a fixed teaching fixture.
Autoregressive and diffusion generation are different loops
Put the two side by side:
Autoregressive
known prefix
→ predict distribution for one next token
→ select token
→ append token
→ repeat
Diffusion
random noisy state
→ predict a correction/noise component
→ update the whole current state
→ repeat
Autoregressive language generation grows a sequence one token at a time.
Diffusion usually refines a continuous state over multiple denoising steps.
Both are generative because they can produce new samples from learned patterns, but they organize generation differently.
Conditioning steers the denoising process
Unconditional diffusion asks only:
What data-like sample should this noisy state become?
A conditioned diffusion model also receives information about what kind of sample is wanted.
The condition might be:
- a class label such as “cat”;
- a text representation such as “a red bicycle in snow”;
- another image or structural signal.
Conceptually:
noisy state + condition
↓
denoiser
↓
correction consistent with both
the data distribution and the condition
Modern text-to-image systems often use text-derived conditioning, and many operate in a learned latent representation rather than directly denoising full-resolution pixels. Those engineering choices matter, but they are not required to understand the generative loop.
A diffusion model is not searching a database for a matching training image
The generation mechanism described above does not perform:
prompt → search training image database → return closest image
Instead, training adjusts model parameters so the denoiser learns statistical structure useful for removing noise. At generation time, the model repeatedly computes corrections from its current state and condition.
That does not mean memorization is impossible. Any high-capacity model can memorize or reproduce training examples under some conditions, and responsible evaluation should test for that risk.
The important distinction is about the mechanism: diffusion generation is learned iterative denoising, not ordinary nearest-image database lookup.
Where GANs and VAEs fit
Diffusion is not the only non-autoregressive generative family.
Two important earlier families are:
- Variational autoencoders (VAEs): learn a structured latent space and decode samples from it;
- Generative adversarial networks (GANs): train a generator against a discriminator that tries to distinguish generated from real examples.
They matter historically and technically, but this Level does not turn them into separate core Lessons. The main Level project remains the tiny language model.
Keep the branch connected to the Level 6 spine
This Lesson broadens your map of modern generative AI. It does not replace the decoder-only language-model path.
After this Lesson, Level 6 returns to the tiny-LLM workflow:
tiny LLM training
→ sampling and decoding
→ brief diffusion branch
→ generation failure diagnosis
→ evaluation
→ packaging
→ integration project
Quick Check
Explain it back
Explain diffusion to someone who knows next-token generation but has never seen an image generator.
Use this order:
- clean data;
- controlled noising;
- learn to predict or remove noise;
- start from noise and denoise repeatedly;
- add conditioning if you want to steer the result.
Then compare that loop with autoregressive token generation in two or three sentences.
Key Takeaways
- Diffusion is a distinct modern generative paradigm built around controlled noising and learned denoising.
- Training creates corrupted examples and teaches a model to predict a useful reverse correction, often the added noise.
- Generation starts from noise and repeatedly refines the current state.
- Conditioning supplies extra information that can steer the denoising trajectory.
- The mechanism is learned iterative denoising, not ordinary database lookup.
- This toy Lab isolates the process; it is not equivalent to training a production text-to-image system.
Next Lesson
Next, return to the tiny-language-model spine in L6.11 — Generation Failure Modes and diagnose repetition, collapse, incoherence, and prompt drift with controlled evidence.
References
- Ho, Jain, and Abbeel, Denoising Diffusion Probabilistic Models.
- Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models.
Completion is stored locally on this device.