본문으로 건너뛰기
L4.7

Token IDs and Special Tokens

Goal

Reserve stable IDs for <pad>, <unk>, <bos>, and <eos>, and explain why changing those mappings can corrupt previously encoded data even when every integer remains valid.

Ordinary pieces represent text. Special tokens represent control information.

<pad> filler position
<unk> piece not represented by the vocabulary
<bos> sequence begins
<eos> sequence ends

Suppose an encoded example is [2, 8, 11, 3]. It is only meaningful if every component that reads it agrees that ID 2 means <bos> and ID 3 means <eos>. If another tokenizer version swaps those meanings, the array is still syntactically valid but semantically wrong.

IDs are conventions that must stay synchronized​

Suppose a tokenizer uses:

<pad> = 0
<bos> = 1
<eos> = 2
cat = 17
sat = 42

Then the sequence:

<bos> cat sat <eos>

becomes:

[1, 17, 42, 2]

Those numbers only have meaning relative to this exact vocabulary.

If another tokenizer says 17 = dog, loading the same model weights with that tokenizer silently changes the input meaning. Shapes still match, but the model receives the wrong symbols.

Special tokens are part of the model contract​

Special tokens can mark boundaries or carry control information:

  • <bos> — beginning of sequence;
  • <eos> — end of sequence;
  • <pad> — fill unused positions in a batch;
  • task- or chat-specific control tokens — distinguish roles or sections.

Their IDs must be consistent across training, evaluation, saving, and generation.

Do not treat special tokens as decorative strings. They become ordinary integer IDs at model input, so a mismatch is a real data bug.

Adding one token can shift later assumptions​

Suppose a sequence originally starts at position 0 with cat. After inserting <bos>, cat moves to position 1.

The token ID for cat may stay 17, but its position changes. That is why token identity and position identity must be handled separately in later lessons.

Predict

What is the safest way to keep control-token meanings stable while ordinary vocabulary entries change?

Verify the contract​

The Lab reserves the four special tokens first, then adds the sorted ordinary words from TEXTS.

  1. Click Run. Read special ids: {'<pad>': 0, '<unk>': 1, '<bos>': 2, '<eos>': 3} and encoded: [2, 4, 1, 3] for "cats quokka": beginning marker, cats, unknown quokka, end marker.
  2. Add a new training sentence. Change TEXTS = ["cats nap", "dogs play"] to TEXTS = ["cats nap", "dogs play", "birds sing"].
  3. Before running, predict: which IDs will stay the same, and will the ID for cats change?
  4. Click Run. The special IDs are still 0–3, but the encoding becomes [2, 5, 1, 3]: cats moved from ID 4 to 5, because birds sorts before it. Reserved IDs protect the control tokens; ordinary IDs can still shift when the vocabulary is rebuilt, which is why a saved vocabulary must travel with saved data.
  5. Press Reset.

Loading lab…

  1. Now create the dangerous version. Change SPECIAL_TOKENS = ["<pad>", "<unk>", "<bos>", "<eos>"] to SPECIAL_TOKENS = ["<eos>", "<unk>", "<bos>", "<pad>"] and run again.
  2. The new encoding is [2, 4, 1, 0]. Now read the old encoding [2, 4, 1, 3] with the new table: its final 3 now means <pad>, so the sentence appears to have no end marker. Nothing crashed and every shape still matches. The bug is only in the meaning.
  3. Press Reset afterward.

Quick Check

1. Why is `<eos>` different from `<pad>`?
2. What can happen if old IDs are read with a new incompatible mapping?
3. Which artifacts must agree on special IDs?

0 of 3 questions answered.

Explain it back​

Explain why a special-token mismatch can look like a model-quality failure even when the model parameters themselves are unchanged.

Key Takeaways

  • Special tokens carry control semantics, not ordinary text meaning.
  • Their IDs are a versioned compatibility contract.
  • Semantic mapping bugs can survive all shape/type checks.
  • Verify configured IDs against the artifact that produced encoded data.

Next Lesson

Complete the checkpoint now. After it, you will use <pad> together with masks and truncation to create rectangular model inputs without confusing filler with real text.

References

Lesson actions

Completion is stored locally on this device.

View progress