Byte-Level Tokenization
Goal
Encode Unicode text as UTF-8 bytes, explain the coverage benefit of a 256-value byte vocabulary, and identify the sequence-length and boundary costs of multi-byte characters.
Bytes give a strong coverage guarantee
The earlier tokenizer lesson showed a simple fixed vocabulary. Now consider text the vocabulary never anticipated: é, 🙂, or another writing system.
UTF-8 converts text into byte values from 0 to 255. That gives a powerful guarantee: any valid UTF-8 string can be represented without inventing an unknown character token.
The guarantee has a cost. One visible character is not always one byte. ASCII A uses one byte, while many accented characters and emoji use multiple bytes, so visible text can expand into several sequence positions.
Follow one character through the layers
Suppose an unfamiliar character becomes four UTF-8 bytes.
The tokenizer may expose those bytes directly or use learned byte-derived subwords. The model then receives token IDs corresponding to those pieces. Later embedding layers turn the IDs into vectors.
The model is not “reading raw Unicode” inside attention. The tokenizer has already converted the text to a finite sequence of discrete IDs.
Coverage is not the same as efficiency
A tokenizer that can represent every input is not automatically a good tokenizer.
If ordinary text regularly expands into very long byte sequences, the model spends more context-window space and more computation on the same human-visible sentence.
That is why practical systems often combine broad byte coverage with learned merges or subwords. Common patterns become compact tokens, while byte-level behavior remains a fallback for unusual text.
When evaluating tokenizers, check both can this text be represented? and how many tokens does it require?
Predict
Inspect actual bytes
Before the lab, compare len(text) with len(text.encode('utf-8')) for A, é, and an emoji. The difference makes the coverage/length tradeoff concrete.
- Click Run. Compare
characters=withbytes=:Ais 1 byte,éis 2, and the emoji🙂is 4. The last line,round trip: Token🙂, shows that bytes decode back to the original text exactly. - Add a character from another writing system. Change
SAMPLES = ["A", "é", "🙂", "AI🙂"]toSAMPLES = ["A", "é", "🙂", "AI🙂", "가"]. - Before running, predict how many bytes the Korean syllable
가needs. - Click Run.
가is 1 character but 3 bytes:[234, 176, 128]. Byte-level tokenization can represent any text, but some scripts cost more positions.
Loading lab…
- Now create a deliberate failure by cutting through the middle of a multi-byte character. Change
round_trip = decode_bytes(byte_tokens("Token🙂"))toround_trip = decode_bytes(byte_tokens("Token🙂")[:-2]), which drops the last two of the emoji's four bytes. - Click Run. The Lab stops with
UnicodeDecodeError ... unexpected end of data. The remaining bytes are not valid UTF-8. This matters later when truncation or windowing cuts a sequence near such a boundary. - Press Reset afterward.
Quick Check
Transfer the idea
For user-generated text with names, symbols, and many languages, explain why broad coverage may be worth extra sequence positions—and why you would still measure token counts rather than assuming the cost is small.
Key Takeaways
- UTF-8 gives a fixed byte alphabet that can represent any valid text.
- Byte coverage avoids unknown characters but can increase sequence length.
- Visible-character boundaries and byte boundaries are different.
- Round-trip and boundary tests are essential evidence for byte representations.
Next Lesson
Next, use corpus statistics to combine frequent smaller pieces into reusable subwords that can reduce sequence length while keeping fallback coverage.
References
- Kudo and Richardson, SentencePiece.
- Radford et al., Language Models are Unsupervised Multitask Learners.
Completion is stored locally on this device.