Skip to main content
L6.11

Generation Failure Modes

Goal

Describe generation failures precisely and use controlled evidence to distinguish plausible data, training, context, tokenizer, and decoding causes.

“Bad output” is a symptom category that is too broad to debug. Start by naming what you observe:

  • repetition — loops or repeated phrases;
  • collapse — many prompts receive nearly the same continuation;
  • incoherence — local token plausibility without stable structure;
  • prompt drift — the continuation rapidly loses the prefix's topic or pattern.

A tiny model can show several at once, and no symptom proves a unique cause.

Turn one vague complaint into testable observations​

Suppose a sample reads:

the cat sat and sat and sat and sat ...

“Bad generation” does not tell you what to measure. “The phrase and sat repeats four times within 12 generated tokens” does.

Now you can form competing hypotheses:

  • duplicated training text may have made the loop unusually likely;
  • the checkpoint may be overfit;
  • low temperature or a narrow candidate filter may be making the distribution too concentrated;
  • context truncation may have removed information that would normally break the pattern.

Each hypothesis suggests a different controlled check. For example, changing only temperature can test a decoding hypothesis, while inspecting train/validation curves tests an overfitting hypothesis.

The important habit is to avoid jumping from one visible symptom directly to one favored cause. Generation is the end of a long chain, so several upstream problems can look similar at the output.

The same visible symptom can come from different layers of the system​

Suppose generation repeats:

blue blue blue blue ...

Possible causes include:

  • the checkpoint strongly prefers a narrow local continuation;
  • decoding settings concentrate probability too aggressively;
  • training data contains repeated patterns;
  • the prompt/context makes a loop unusually likely;
  • the tokenizer creates an awkward representation around the repeated text.

That is why symptom → single cause is unsafe reasoning.

A better workflow is:

  1. describe the failure precisely;
  2. keep the checkpoint fixed and test decoding hypotheses;
  3. inspect whether the failure appears across several prompts;
  4. compare nearby checkpoints;
  5. inspect relevant data or tokenization only when evidence points upstream.

Each test should eliminate or strengthen one hypothesis.

Build fixed failure cases​

If a particular prompt exposes a repeat loop, save it as an evaluation case with the exact checkpoint and decoding configuration.

After changing the system, rerun the same case.

Otherwise “it seems better now” may simply mean you happened to sample a different prompt or random outcome.

Predict

A model emits nearly the same short phrase for many distinct prompts under fixed moderate sampling. What is the best initial description?

Use an evidence ladder​

The Lab measures one simple failure indicator: the repeated-bigram rate, the share of neighboring word pairs that already appeared earlier in the same text.

  1. Click Run. The looping sample scores loop repetition : 0.714; the varied sample scores 0.0.
  2. Introduce a known repetition. Change varied = "the cat sat near a warm window" to varied = "the cat sat near the cat sat".
  3. Before running, predict whether the varied score will rise, and whether it will reach the loop's score.
  4. Click Run. varied repetition rises to 0.333: two of its six word pairs are repeats. The measurement moved in the expected direction, but it stays below the obvious loop.
  5. Press Reset afterward.

Loading lab…

For a real failure, keep checkpoint, prompt, seed, and decoding config fixed, then inspect in order:

  1. train/validation evidence;
  2. representative training text and duplication;
  3. one decoding control at a time;
  4. context truncation/windowing;
  5. tokenizer compatibility/segmentation;
  6. only then architecture or capacity changes.

This order does not claim every failure has a simple cause. It prevents expensive changes from hiding basic evidence.

Quick Check

1. What is the best first move when generation looks wrong?
2. Can repetition arise from either data or decoding choices?
3. Why is incoherence in a tiny model not automatically a software bug?

0 of 3 questions answered.

Explain it back​

For a repeated phrase loop, write three hypotheses and one observation that would strengthen or weaken each. Keep “symptom” separate from “cause.”

Key Takeaways

  • Name observable generation failures before naming causes.
  • Reproduce a failure under fixed checkpoint/prompt/decoding conditions.
  • Use training, data, context, tokenizer, and decoding evidence to separate hypotheses.
  • Some poor output reflects genuine tiny-model limits rather than broken code.

Next Lesson

Next, combine held-out prediction evidence and fixed generation cases into an evaluation that is more reliable than one favorite sample.

References

Lesson actions

Completion is stored locally on this device.

View progress