본문으로 건너뛰기
L4.13

Tokenizer Failure Modes

Goal

Diagnose tokenizer failures using concrete text slices, unknown rate, token count/inflation, round-trip checks, and artifact compatibility evidence.

A tokenizer can produce valid arrays for every input and still be a bad representation. Three common patterns are:

  • coverage loss — meaningful distinctions collapse to <unk>;
  • token inflation — ordinary text consumes many more positions than expected;
  • contract mismatch — special IDs, normalization, or vocabulary versions disagree across artifacts.

These failures matter because they happen before the language model gets a chance to reason about the text.

Diagnose the symptom before changing the tokenizer​

Tokenizer problems can appear in several different ways.

Unknown or fallback explosion: important text collapses into unknown/fallback pieces.

Sequence inflation: ordinary-looking text requires unexpectedly many tokens.

Boundary instability: small text changes produce surprising segmentation changes that matter downstream.

Domain mismatch: code, equations, names, or another language are represented much less efficiently than the tokenizer's training domain.

These symptoms have different causes, so “the tokenizer is bad” is not a useful diagnosis by itself.

Compare examples from the real domain​

Suppose an English paragraph uses 80 tokens while an equally sized code snippet uses 210.

That observation does not prove the tokenizer is unusable. It tells you to measure whether code-token inflation creates practical costs:

  • fewer source lines fit in the context window;
  • training and inference process more token positions;
  • important structures may be split into many small pieces.

A fair tokenizer evaluation therefore uses representative samples rather than one convenient sentence.

Do not blame tokenization for every language-model failure​

If a model generates a wrong factual answer but the prompt was encoded cleanly and compactly, the tokenizer may not be the cause.

Conversely, a tokenizer can encode every character without unknown tokens and still produce inefficient sequences.

Separate at least three questions:

  1. coverage: can the text be represented?
  2. efficiency: how many tokens does it require?
  3. model behavior: what does the trained model do with those tokens?

Keeping these layers separate prevents random tokenizer changes from becoming a substitute for actual failure analysis.

Predict

Which metric directly shows how often a word-level tokenizer falls back to `<unk>`?

Test slices, not anecdotes​

The Lab compares a word-level tokenizer (trained on two short sentences) with byte-level encoding on four fixed cases in EVAL_CASES.

  1. Before running, predict for each case which representation will struggle: familiar, rare_name, emoji, and multilingual.
  2. Click Run. The word tokenizer handles familiar perfectly (word_unknown_rate: 0.0) but loses everything in rare_name and multilingual (1.0). Byte encoding never produces an unknown piece, but it needs many more positions, for example 19 bytes for three familiar words.
  3. Add a code/number slice. In EVAL_CASES, add a new line: "code": "x = 42",.
  4. Click Run. The new line shows word_unknown_rate: 1.0 with only 6 bytes: the word vocabulary has never seen x, =, or 42, while bytes cover it cheaply.
  5. Press Reset afterward.

Loading lab…

Keep the exact produced pieces for the worst examples. An aggregate average can hide a severe failure concentrated in one group of text.

Quick Check

1. Why is zero runtime errors weak tokenizer evidence?
2. What does token inflation describe?
3. Why preserve exact failing strings and pieces?

0 of 3 questions answered.

Explain it back​

Explain one way tokenizer quality could differ across user groups even if downstream model architecture and weights are identical.

Key Takeaways

  • Tokenizer failures include information loss, inefficiency, and compatibility errors.
  • Measure representative slices beyond the tokenizer's training examples.
  • Aggregate metrics need concrete worst-case examples beside them.
  • Representation problems should be fixed before blaming downstream model reasoning.

Next Lesson

Next, combine these measurements in a controlled tokenizer evaluation and make an evidence-based engineering choice.

References

Lesson actions

Completion is stored locally on this device.

View progress