본문으로 건너뛰기
L6.10

Diffusion Models: Learning to Denoise

Goal

By the end of this lesson, you can explain diffusion generation as learned iterative denoising and compare its generation loop with autoregressive next-token generation.

Level 6 has followed one modern generative pattern so far:

Autoregressive generation
token → next token → next token → ...

Now consider a different job.

Suppose you have a clean low-resolution pattern:

█ · █ ·
· █ · █
█ · █ ·
· █ · █

Add a little random corruption and the structure is still visible. Add more and more corruption and eventually the original pattern is hard to recognize.

A diffusion model learns the opposite skill: given a noisy state, predict how to remove some of the noise.

Generation then runs that learned repair process repeatedly:

Diffusion generation
noise → less noise → less noise → ... → sample

The important idea comes before any architecture details.

Step 1: start with clean data​

Training begins with ordinary examples from the data distribution: images, audio representations, or other continuous data.

Call one clean example x₀.

At this point there is no generation yet. We already have the clean training example.

Step 2: add noise in controlled steps​

Imagine corruption levels from 0 to 4:

step 0 clean
step 1 slightly noisy
step 2 noisier
step 3 heavily noisy
step 4 almost entirely noise

The forward noising process is deliberately easy to run. We control how much noise is added at each step.

This process is not the generative direction. It creates training situations where the model can practice recognizing what corruption was added.

Predict

Why deliberately corrupt clean training examples?

Step 3: train a model to predict what should be removed​

For a noisy example, the model receives information such as:

  • the corrupted sample;
  • how far along the noising process it is;
  • optionally, a conditioning signal.

A common training target is the noise that was added.

clean example
↓ add known noise
noisy example
↓ denoiser
predicted noise
↓ compare
actual added noise

If the predicted noise is wrong, the loss is larger. Optimizer updates gradually improve the denoiser.

Only now, after seeing the corruption-and-repair job, do we need a small amount of math. A toy corruption rule can be written as:

corrupted = (1 - noise_level) × clean + noise_level × noise

This linear blend is only a teaching rule for making the corruption visible. It is not the exact forward equation used by DDPMs. Real diffusion systems usually draw random noise from a defined distribution and combine it with the clean example using timestep-dependent coefficients.

A simple training loss can compare predicted noise with the known noise:

loss = mean((predicted_noise - added_noise)²)

Real diffusion systems use more carefully designed noise schedules and parameterizations. You do not need scheduler notation, stochastic differential equations, or a particular neural-network architecture to understand the core training signal.

Step 4: generate by starting from noise​

At generation time there is no clean answer to reveal.

Instead:

  1. start from random noise;
  2. ask the trained denoiser what noise-like component should be removed;
  3. move to a slightly cleaner state;
  4. repeat for several steps.

That is why sampling is iterative.

One denoising step usually does not produce a final image-like sample any more than one next-token decision produces a whole paragraph.

Run a visible corruption-and-repair experiment​

The Lab uses a four-number pattern instead of an image so every value remains inspectable.

It contains:

  • one clean pattern;
  • one fixed noise vector, standing in for one reproducible random draw;
  • five forward corruption levels;
  • a precomputed sequence of idealized reverse corrections;
  • a learner-controlled denoise_strength.

The precomputed corrections are deliberately perfect teaching fixtures: each one points from the current toy state toward a cleaner state. They are written directly into the fixture rather than calculated from the clean target during reverse generation. The clean pattern is used afterward only to measure error. These corrections are not an implementation of a DDPM's learned noise-prediction network. This Lab isolates the reverse-process intuition and does not claim to train a production image model in the browser.

Loading lab…

Use this sequence:

  1. Click Run unchanged.
  2. Read the forward errors from clean. They should grow as the noise level increases.
  3. Read the reverse errors from clean. They should shrink as denoising proceeds.
  4. Confirm that the reverse process starts from the fixed noise vector, which stands in for one reproducible random start, not from the clean target.
  5. Change only denoise_strength = 1.0 to 0.5.
  6. Predict what will happen: each correction is smaller, so after the same number of steps the final state should remain farther from the clean pattern.
  7. Run and compare the final error.
  8. Reset.

This is a process experiment. Notice the important boundary: run_reverse starts from the fixed noise vector and consumes only the prewritten corrections plus denoise_strength; it does not consult clean. A real diffusion model would produce those updates with a trained denoiser rather than a fixed teaching fixture.

Autoregressive and diffusion generation are different loops​

Put the two side by side:

Autoregressive
known prefix
→ predict distribution for one next token
→ select token
→ append token
→ repeat

Diffusion
random noisy state
→ predict a correction/noise component
→ update the whole current state
→ repeat

Autoregressive language generation grows a sequence one token at a time.

Diffusion usually refines a continuous state over multiple denoising steps.

Both are generative because they can produce new samples from learned patterns, but they organize generation differently.

Conditioning steers the denoising process​

Unconditional diffusion asks only:

What data-like sample should this noisy state become?

A conditioned diffusion model also receives information about what kind of sample is wanted.

The condition might be:

  • a class label such as “cat”;
  • a text representation such as “a red bicycle in snow”;
  • another image or structural signal.

Conceptually:

noisy state + condition
↓
denoiser
↓
correction consistent with both
the data distribution and the condition

Modern text-to-image systems often use text-derived conditioning, and many operate in a learned latent representation rather than directly denoising full-resolution pixels. Those engineering choices matter, but they are not required to understand the generative loop.

A diffusion model is not searching a database for a matching training image​

The generation mechanism described above does not perform:

prompt → search training image database → return closest image

Instead, training adjusts model parameters so the denoiser learns statistical structure useful for removing noise. At generation time, the model repeatedly computes corrections from its current state and condition.

That does not mean memorization is impossible. Any high-capacity model can memorize or reproduce training examples under some conditions, and responsible evaluation should test for that risk.

The important distinction is about the mechanism: diffusion generation is learned iterative denoising, not ordinary nearest-image database lookup.

Where GANs and VAEs fit​

Diffusion is not the only non-autoregressive generative family.

Two important earlier families are:

  • Variational autoencoders (VAEs): learn a structured latent space and decode samples from it;
  • Generative adversarial networks (GANs): train a generator against a discriminator that tries to distinguish generated from real examples.

They matter historically and technically, but this Level does not turn them into separate core Lessons. The main Level project remains the tiny language model.

Keep the branch connected to the Level 6 spine​

This Lesson broadens your map of modern generative AI. It does not replace the decoder-only language-model path.

After this Lesson, Level 6 returns to the tiny-LLM workflow:

tiny LLM training
→ sampling and decoding
→ brief diffusion branch
→ generation failure diagnosis
→ evaluation
→ packaging
→ integration project

Quick Check

1. What is the purpose of the forward noising process during diffusion training?
2. How does diffusion generation usually begin?
3. What does conditioning add conceptually?
4. Which comparison is accurate?

0 of 4 questions answered.

Explain it back​

Explain diffusion to someone who knows next-token generation but has never seen an image generator.

Use this order:

  1. clean data;
  2. controlled noising;
  3. learn to predict or remove noise;
  4. start from noise and denoise repeatedly;
  5. add conditioning if you want to steer the result.

Then compare that loop with autoregressive token generation in two or three sentences.

Key Takeaways

  • Diffusion is a distinct modern generative paradigm built around controlled noising and learned denoising.
  • Training creates corrupted examples and teaches a model to predict a useful reverse correction, often the added noise.
  • Generation starts from noise and repeatedly refines the current state.
  • Conditioning supplies extra information that can steer the denoising trajectory.
  • The mechanism is learned iterative denoising, not ordinary database lookup.
  • This toy Lab isolates the process; it is not equivalent to training a production text-to-image system.

Next Lesson

Next, return to the tiny-language-model spine in L6.11 — Generation Failure Modes and diagnose repetition, collapse, incoherence, and prompt drift with controlled evidence.

References

Lesson actions

Completion is stored locally on this device.

View progress