Skip to main content
L6.9

Sampling Text

Goal

Explain generation as a repeated next-token selection loop and use a fixed probability distribution plus seed to distinguish stochastic sampling from model training.

A language model produces one next-token distribution at a time​

Training changes model parameters. Sampling does not. Sampling starts after the model has produced scores for the next token.

A decoder generates text by repeating this loop:

current context
↓ model
next-token distribution
↓ choose one token
append that token
↓
new context
↺

For example:

"the moon"
↓ model
P(next token)
↓ choose "is"
"the moon is"
↓ model
P(next token)
↓ choose "quiet"
"the moon is quiet"

The newly selected token becomes part of the next input. That is what autoregressive generation means here: each new choice is conditioned on the prompt plus the tokens already generated.

Predict

A next-token distribution gives A=.6, B=.3, C=.1. Does sampling have to choose A every time?

Make randomness reproducible before changing policy​

The Lab deliberately uses a tiny fixed distribution:

A: 0.6
B: 0.3
C: 0.1
seed: 7

This is smaller than a real vocabulary so you can see the sampling idea without mixing it with temperature, top-k, or top-p yet.

  1. Click Run unchanged.
  2. Read the 12 sampled tokens: A A B A A A A A A A A A. Mostly A, as the weights suggest, but not only A.
  3. Click Run again without changing the weights or the seed.
  4. Confirm that the 12-token sample repeats exactly.
  5. Change only rng = random.Random(7) to rng = random.Random(8), then run again.
  6. The draws become A C A B A A C A B A A A: a different sequence, even though the token weights stay [0.6, 0.3, 0.1].
  7. Press Reset afterward.

Loading lab…

The evidence separates two things:

  • the distribution says which outcomes are more or less likely;
  • the seed makes a particular sequence of random draws reproducible for an experiment.

A seed does not make A more probable than B. The weights already define that relationship.

Sampling is not greedy decoding​

If you always choose the token with the largest score, that is greedy decoding. It is deterministic for fixed logits.

Sampling instead draws from a probability distribution, so several continuations can be possible from the same model state.

Neither method makes the underlying model better trained. They are different ways to choose from what the model already predicted.

Debug the right layer​

If generated text looks bad, first separate these questions:

  1. Model question: Did training and validation produce a useful next-token distribution?
  2. Sampling question: Given that distribution, are we selecting tokens in the way we intended?

A decoding change cannot repair a model that never learned useful patterns. Conversely, a good checkpoint can still produce different text when the random draws or later decoding controls change.

Quick Check

1. Why append each selected token to the context?
2. What does a fixed sampling seed give you?
3. What changes when you sample another token from fixed model logits?

0 of 3 questions answered.

Key Takeaways

  • Generation repeats next-token prediction, selection, append, and predict again.
  • Sampling can choose lower-probability tokens; it is not the same as greedy argmax.
  • A fixed seed makes stochastic experiments reproducible under fixed conditions.
  • Sampling changes the selected continuation, not the trained model weights.
  • Diagnose model quality separately from decoding behavior.

Next Lesson

Next, L6.10 — Temperature, Top-k, and Top-p keeps the sampling loop fixed and changes the decoding policy: which next-token choices remain likely or eligible.

References

Lesson actions

Completion is stored locally on this device.

View progress