Build a Simple Tokenizer
Goal
Turn text into stable integer token IDs, then turn those IDs back into readable tokens, while making the limitations of a tiny word-level vocabulary visible.
A language model cannot multiply words directly. It needs numbers. Tokenization is the reproducible bridge from text to the integer IDs that later select embedding rows.
Entry skills for this coding lesson
You should be comfortable reading basic Python variables, lists, dictionaries, functions, and for loops. You do not need to know regular expressions, tensors, matrix multiplication, softmax, PyTorch modules, gradients, or cross-entropy yet.
If those basic Python pieces are unfamiliar, use the Project Workbench bridge first. Do not treat a Python-environment problem as if it were a tokenizer concept problem.
One sentence, two representations
Start with:
The model predicts the next token.
A simple word-level tokenizer can split it into:
["the", "model", "predicts", "the", "next", "token", "."]
A vocabulary then gives each distinct piece a stable integer ID. The exact numbers are arbitrary; the mapping is not. If "model" is ID 7 today and ID 12 tomorrow while the model weights stay unchanged, the same integer no longer means the same input.
Saved tokenizer data and model weights therefore have to agree on what every ID means.
Predict
Make the splitting rule explicit
This deterministic splitter is deliberately small:
import re
pattern = re.compile(r"[A-Za-z]+|[0-9]+|[^\w\s]")
def split(text):
return pattern.findall(text.lower())
print(split("Tiny models learn, too!"))
# ['tiny', 'models', 'learn', ',', 'too', '!']
A regular expression (regex) is a compact text pattern. Here it means: take runs of letters, runs of digits, or one punctuation symbol. You do not need to design regex syntax for this Level. You only need to know exactly what rule produced the pieces.
Build a stable vocabulary
A sorted vocabulary makes the same corpus produce the same IDs every run:
tokens = split(corpus)
vocab = ["<unk>", *sorted(set(tokens))]
stoi = {token: i for i, token in enumerate(vocab)}
itos = vocab
def encode(text):
return [stoi.get(token, 0) for token in split(text)]
def decode_tokens(ids):
return [itos[i] for i in ids]
<unk> uses ID 0. Any piece missing from the vocabulary falls back to that explicit unknown token rather than receiving a random new ID.
Run the tokenizer notebook
The Lab opens in a notebook and runs on CPU; you do not need a GPU.
- Open the Lab and run the cells unchanged.
- In the sample cell, find:
sample = "The model predicts the next token."
- Read the printed token list, ID list, and decoded text. Confirm that the number of tokens equals the number of IDs.
- In the later failure cell, run
tokenizer.encode("the zzzxxyy token")and inspect bothunknown_idsanddecode_tokens(unknown_ids). - Find the
0inunknown_idsand the matching<unk>piece. That is visible evidence of information the tiny vocabulary cannot represent exactly. - Do not “fix” the example by silently adding
zzzxxyyto the vocabulary. The limitation is the point of this experiment.
Loading lab…
After the guided pass, change only the sample string to another sentence made from familiar corpus words and punctuation. Predict its pieces before rerunning that cell.
What can go wrong even when the code runs?
If two machines build different IDs from the same text, check these boundaries first:
- normalization — did both lowercase the same way?
- splitting — did both produce the same pieces?
- vocabulary ordering — did both assign IDs in the same deterministic order?
- saved tokenizer identity — are both actually using the same saved tokenizer and vocabulary?
A set can collect unique pieces, but sorting those pieces makes the teaching mapping explicit and reproducible.
Why real LLM tokenizers go beyond whole words
A word-level tokenizer is easy to inspect, but unseen words collapse to <unk>. Real LLMs commonly use subword or byte-oriented strategies so they can represent unfamiliar names, spellings, languages, and symbols without requiring one vocabulary entry for every possible word.
You will explore those tradeoffs next rather than hiding them inside this first tokenizer.
Quick Check
Key Takeaways
- Tokenization converts text into discrete pieces and stable integer IDs.
- ID values are arbitrary addresses; the tokenizer mapping gives them meaning.
- Deterministic splitting and vocabulary ordering make runs reproducible.
<unk>makes an unseen-word limitation visible instead of hiding it.- Whole-word tokenization is intentionally simple; the next lessons explore byte and subword alternatives.
Next Lesson
Next, you will study L4.5 — Byte-level Tokenization and see how UTF-8 bytes can represent any input text without a whole-word <unk> fallback, at the cost of potentially longer sequences.
References
- PyTorch embedding documentation: https://pytorch.org/docs/stable/generated/torch.nn.Embedding.html
- SentencePiece paper: https://arxiv.org/abs/1808.06226
Completion is stored locally on this device.