본문으로 건너뛰기

Level 4 Checkpoint — Represent Text with Stable Tokens

Complete this checkpoint after L4.7 — Token IDs and Special Tokens. The goal is not to recite tokenizer names; it is to explain what information each representation keeps, loses, or makes more expensive.

1. One string, three granularities​

Use players 🙂.

Represent it approximately as whole-word pieces, reusable subwords, and UTF-8 bytes. For each representation, state:

  • what gives it coverage;
  • what can make its sequence longer;
  • where an unknown fallback could occur.

2. Unknown-token information loss​

A word vocabulary maps both quokka and axolotl to <unk>.

Explain precisely what the downstream model can no longer distinguish at that token position. Then give one representation strategy that could preserve more of the spelling.

3. Subword merge reasoning​

A frequent pair appears 12 times in a corpus. If it is merged, explain why the vocabulary can become larger while those 12 occurrences become shorter. Then name one reason the merge might not help on a different domain.

4. Stable ID rule​

Reserve:

0 <pad>
1 <unk>
2 <bos>
3 <eos>

Add three ordinary pieces without changing those IDs. Then explain what can go wrong if a later tokenizer version swaps IDs 0 and 3 while old encoded data is reused.

The important rule is simple: saved data and the model must agree on what each special ID means. Engineers sometimes call a must-agree rule like this a contract.

5. Check the rule in the Lab​

Run the Lab below. Add an unfamiliar word, verify its reserved unknown ID, and confirm that the beginning/end IDs remain unchanged.

Loading lab…

Quick Check

1. What does a token ID mean without its vocabulary?
2. Why can byte tokenization represent unfamiliar Unicode text?
3. What is the direct merge tradeoff?
4. Why must special IDs stay stable?

0 of 4 questions answered.

Check your reasoning after you try

A useful answer should make the tradeoffs concrete:

  • Whole-word tokenization is short for known words but needs a fallback for unseen words such as rare names.
  • Subwords can spell unfamiliar words from reusable pieces, usually with more positions than a known whole word.
  • UTF-8 bytes always have fixed coverage through byte values 0–255, but one visible character can require several positions.
  • Mapping both quokka and axolotl to <unk> erases which spelling occurred at that position.
  • Stable special-token IDs are a compatibility contract. Reusing old encoded data after swapping <pad> and <eos> can silently change control meaning.

Pass condition​

You are ready for L4.8 — Padding, Truncation, and Masks when you can trace text → pieces → IDs, describe at least one coverage/length tradeoff, and explain a special-ID mismatch in ordinary language without relying on “the tokenizer library handles it.”

Lesson actions

Completion is stored locally on this device.

View progress