Temperature, Top-k, and Top-p
Goal
Explain how temperature, top-k, and top-p change one next-token choice, calculate their effect on a small distribution, and compare decoding settings without confusing decoding changes with model training.
L6.9 — Sampling Text separated model prediction from token selection. The model produces logits. A decoding policy decides how those logits are turned into an eligible probability distribution and, finally, a selected token.
The trained weights do not change during this process.
Start with one fixed logit vector
Suppose the model produces these four logits:
A: 2.0
B: 1.0
C: 0.2
D: -0.5
Softmax turns them into probabilities. The exact values are less important than the ordering:
A > B > C > D
All three controls in this lesson act on this one next-token decision. They answer different questions:
- temperature: how sharp or flat should the relative preferences be?
- top-k: how many of the highest-scoring candidates may remain?
- top-p: how much cumulative probability mass should the retained candidate set cover?
None of these asks the model to learn new weights.
Temperature changes score gaps before softmax
Temperature is applied to logits:
softmax(logits / temperature)
For a positive temperature below 1, dividing by a small number makes logit gaps larger. Softmax then becomes more concentrated on the highest logits.
For a temperature above 1, the gaps shrink and the distribution becomes flatter.
Consider just two logits, [2, 1]:
temperature = 0.5 → scaled logits [4, 2] → sharper
temperature = 1.0 → scaled logits [2, 1] → baseline
temperature = 2.0 → scaled logits [1, 0.5] → flatter
The ordering does not change because every logit is divided by the same positive number.
A common mistake is to divide already normalized probabilities by temperature and call it the same operation. It is not. Temperature belongs on logits before softmax.
Predict
Top-k keeps a fixed number of candidates
Suppose softmax gives:
A=.55 B=.25 C=.12 D=.06 E=.02
With top-k=2, only A and B remain eligible:
before: A=.55 B=.25 C=.12 D=.06 E=.02
keep: A B
Their original probabilities sum to .80, so they are renormalized:
A = .55 / .80 = .6875
B = .25 / .80 = .3125
The retained probabilities must again sum to 1.
top-k=1 leaves only the highest-scoring candidate. After renormalization that candidate has probability 1, so the step behaves greedily.
Top-k does not mean “keep every token above probability k.” The k is a candidate count.
Top-p keeps enough candidates to reach a probability mass
Top-p, often called nucleus sampling, uses a different rule.
Sort candidates from highest to lowest probability, then keep the smallest prefix whose cumulative mass reaches at least p.
Using the same distribution:
A=.55 cumulative=.55
B=.25 cumulative=.80
C=.12 cumulative=.92
...
For top-p=.80, A and B are enough. For top-p=.85, A+B is only .80, so C must also be included.
That threshold-crossing token matters. Stopping before it would fail to reach the requested probability mass.
Top-p therefore keeps a variable number of candidates. A very confident model step might need only one or two tokens to reach .9; a flatter step might need many.
Compare top-k and top-p without memorizing slogans
Imagine two next-token distributions.
Distribution 1: [.90, .04, .03, .02, .01]
Distribution 2: [.25, .22, .20, .18, .15]
With top-k=3, both keep exactly three tokens.
With top-p=.80:
- Distribution 1 reaches .80 with the first token alone.
- Distribution 2 needs several tokens.
That is the central difference: top-k fixes candidate count; top-p adapts candidate count to the shape of the current distribution.
Neither policy is automatically “better.” They create different selection behavior, and their usefulness depends on the model, task, and evaluation criteria.
Run the Lab with one controlled change
The Lab uses the exact logits:
[2.0, 1.0, 0.2, -0.5]
and prints:
- the base softmax distribution;
top-k=2;top-p=0.8.
Loading lab…
Use this sequence:
- Click Run once. The
top_pTODO leaves every candidate active (top-p: [0.619, 0.228, 0.102, 0.051]), so one Lab check fails. That failure is expected. - Complete
top_p: sort candidates by probability, keep the smallest prefix whose cumulative mass reachesp, then renormalize the kept probabilities. - Click Run again and confirm every check passes. With
p = 0.8, the first candidate alone (0.619) is not enough, but the first two (0.619 + 0.228 = 0.847) are, sotop-pbecomes[0.731, 0.269, 0.0, 0.0]. - Count the non-zero candidates: two under
top-k=2, and also two undertop-p=0.8. - Change only
top_k(p1, 2)totop_k(p1, 3). Predict the additional candidate first, then run. The third candidate appears:[0.652, 0.24, 0.108, 0.0]. - Restore
k=2, then change onlytop_p(p1, 0.8)totop_p(p1, 0.9). - Predict whether the retained candidate count stays the same or grows, then run. It grows to three, because two candidates only reach
0.847, which is less than0.9. - Restore
p=0.8.
For a temperature experiment, change only the temperature passed to softmax, leaving the logits fixed. That isolates distribution sharpness from candidate filtering.
Why a prettier sample proves less than it seems
Suppose temperature 0.7 produces a sentence you prefer to the sample from temperature 1.0.
That is evidence that this decoding configuration produced a preferred sample under these conditions. It does not show that:
- the checkpoint learned better parameters;
- validation loss improved;
- every prompt benefits from the same temperature;
- the model became more knowledgeable.
For a fair comparison, keep checkpoint, prompt, random seed policy, maximum generation length, and every other decoding control fixed. Change one variable at a time and evaluate more than one lucky output.
Common mistakes
- Temperature ≤ 0: the usual sampling formula requires a positive temperature; use an explicit greedy rule rather than treating zero as an ordinary temperature.
- Masking after selection: illegal tokens or policy constraints should be handled at the appropriate candidate/logit stage, not after an invalid token has already been sampled.
- Forgetting renormalization: after top-k or top-p removes candidates, the remaining probabilities need to sum to 1 before sampling.
- Changing temperature, k, p, seed, and prompt together: the new text may differ, but you cannot attribute the difference to one cause.
Quick Check
Explain it back
Explain all three controls without using “more random” as your only description. State where each control acts, what it retains or rescales, and what stays unchanged about the trained model.
Key Takeaways
- Temperature rescales logits before softmax and changes distribution sharpness.
- Top-k keeps a fixed number of highest-scoring candidates.
- Top-p keeps the smallest high-probability set whose cumulative mass reaches a threshold.
- Candidate filters must renormalize the retained probabilities before sampling.
- Decoding changes selection behavior; it does not retrain the checkpoint.
- Fair decoding comparisons keep every non-target condition fixed.
Next Lesson
Next, classify bad generations by observable symptom before guessing whether the cause is data, training, context, tokenizer, or decoding.
References
- Holtzman et al., The Curious Case of Neural Text Degeneration.
- PyTorch, torch.topk.
Completion is stored locally on this device.