Level 4 Checkpoint — Represent Text with Stable Tokens
Complete this checkpoint after L4.7 — Token IDs and Special Tokens. The goal is not to recite tokenizer names; it is to explain what information each representation keeps, loses, or makes more expensive.
1. One string, three granularities
Use players 🙂.
Represent it approximately as whole-word pieces, reusable subwords, and UTF-8 bytes. For each representation, state:
- what gives it coverage;
- what can make its sequence longer;
- where an unknown fallback could occur.
2. Unknown-token information loss
A word vocabulary maps both quokka and axolotl to <unk>.
Explain precisely what the downstream model can no longer distinguish at that token position. Then give one representation strategy that could preserve more of the spelling.
3. Subword merge reasoning
A frequent pair appears 12 times in a corpus. If it is merged, explain why the vocabulary can become larger while those 12 occurrences become shorter. Then name one reason the merge might not help on a different domain.
4. Stable ID rule
Reserve:
0 <pad>
1 <unk>
2 <bos>
3 <eos>
Add three ordinary pieces without changing those IDs. Then explain what can go wrong if a later tokenizer version swaps IDs 0 and 3 while old encoded data is reused.
The important rule is simple: saved data and the model must agree on what each special ID means. Engineers sometimes call a must-agree rule like this a contract.
5. Check the rule in the Lab
Run the Lab below. Add an unfamiliar word, verify its reserved unknown ID, and confirm that the beginning/end IDs remain unchanged.
Loading lab…
Quick Check
Check your reasoning after you try
A useful answer should make the tradeoffs concrete:
- Whole-word tokenization is short for known words but needs a fallback for unseen words such as rare names.
- Subwords can spell unfamiliar words from reusable pieces, usually with more positions than a known whole word.
- UTF-8 bytes always have fixed coverage through byte values 0–255, but one visible character can require several positions.
- Mapping both
quokkaandaxolotlto<unk>erases which spelling occurred at that position. - Stable special-token IDs are a compatibility contract. Reusing old encoded data after swapping
<pad>and<eos>can silently change control meaning.
Pass condition
You are ready for L4.8 — Padding, Truncation, and Masks when you can trace text → pieces → IDs, describe at least one coverage/length tradeoff, and explain a special-ID mismatch in ordinary language without relying on “the tokenizer library handles it.”
Completion is stored locally on this device.