본문으로 건너뛰기
L4.3

Vocabulary and Unknown Tokens

Goal

Build a deterministic vocabulary, encode known and unknown pieces, and explain what information is lost when different unseen pieces share one <unk> ID.

Suppose a tiny word vocabulary knows cat, dog, and runs, but not quokka or axolotl. A fallback keeps the program running, but both unseen pieces map to the same reserved entry.

Unknown tokens: different inputs can collapse to one ID

A fallback keeps encoding defined, but it can erase which unseen piece was originally present.

quokka<unk>0
axolotl<unk>0

Information lost: the original unseen pieces are different, but their encoded representation is now identical.

The failure is not a crash. It is information collapse: two different source pieces have become indistinguishable at that position.

A vocabulary therefore has two responsibilities. It must cover useful pieces, and it must map them to IDs deterministically so stored sequences keep the same meaning across runs.

A vocabulary is a finite lookup table​

Imagine a vocabulary containing only:

0: <unk>
1: cat
2: dog
3: runs
4: sleeps

The sentence cat runs can become [1, 3].

But if fox is not representable by smaller known pieces, an old-style word tokenizer may fall back to <unk>. Then fox runs, robot runs, and banana runs can begin with the same ID. That is real information loss.

Coverage and vocabulary size pull in opposite directions​

Putting every possible word in a vocabulary does not scale: language keeps producing names, spelling variants, numbers, URLs, code, and many writing systems. A huge vocabulary also enlarges embedding and output matrices.

Modern tokenizers instead use reusable pieces, bytes, or fallback strategies so rare text can still be represented.

Unknown does not mean semantically unknown​

The token <unk> means the tokenizer could not represent a surface form under its current policy. It does not mean the model has decided the concept is unknown.

Conversely, a tokenizer may encode a rare word perfectly while the model knows very little about its meaning. Tokenizer coverage and model knowledge are different failure layers.

Predict

If two unseen words both become the same `<unk>` ID, what has the representation lost?

Measure coverage instead of assuming it​

Build a tiny vocabulary with <unk> reserved first, then encode one familiar sentence and one from a different topic. The simple diagnostic unknown pieces / total pieces makes coverage failure visible.

  1. Click Run. Read vocab: ['<unk>', 'learn', 'models', 'predict', 'tiny', 'tokens'].
  2. Read evaluation ids: [4, 0, 0] and unknown rate: 0.667. Only tiny is known; quokka and predicts both collapse to ID 0.
  3. Notice predicts. The vocabulary knows predict, but a word-level tokenizer treats predicts as a completely different word.
  4. Add one training sentence. In TRAIN_TEXT, add "a quokka naps", as a third line.
  5. Before running, predict the new unknown rate on the same evaluation sentence.
  6. Click Run. quokka now has its own ID and the rate drops to 0.333, but predicts is still unknown. Coverage improved only for the exact word you added.
  7. Click Run a second time without editing. The vocabulary and IDs are identical, because build_vocab sorts the pieces. That determinism is what keeps stored IDs meaningful across runs.
  8. Press Reset afterward.

Loading lab…

A low unknown rate is not proof that a tokenizer is good; it is one piece of evidence. Byte and subword approaches will show ways to preserve more information while keeping the representation finite.

Quick Check

1. Why is `<unk>` useful?
2. What can a high unknown rate indicate?
3. Why make vocabulary construction deterministic?

0 of 3 questions answered.

Explain it back​

Explain how a tokenizer can have zero runtime errors while still being unusable because its unknown rate is high on the target users' text.

Key Takeaways

  • The vocabulary is the piece-to-ID contract.
  • <unk> is a safe fallback that deliberately loses identity.
  • Coverage must be measured on representative evaluation text.
  • Deterministic vocabulary construction is necessary for reproducibility.

Next Lesson

Next, L4.4 — Build a Simple Tokenizer builds the simple tokenizer end to end. Use it to connect the boundary and vocabulary ideas to runnable code.

References

Lesson actions

Completion is stored locally on this device.

View progress