본문으로 건너뛰기
L4.0

Level 4 — Text Representation and Tokenization

A person can read the sentence:

Cats nap.

A language model cannot start from the visual letters in the way a person does. Before the model can process the sentence, the text must be turned into a reproducible sequence of numbers.

Level 4 follows that conversion step by step.

Start with one sentence​

One possible conversion is:

Cats nap.
→ ["Cats", "nap", "."]
→ [17, 42, 3]
→ three learned vectors of numbers

The pieces such as Cats and nap are tokens in this example. The rule that turns text into tokens and then token IDs is a tokenizer.

The numbers 17, 42, and 3 are token IDs. They are lookup numbers, not measurements of meaning. If a different tokenizer consistently uses 5 for Cats, that can also be valid.

The learned vectors selected by those IDs are called embeddings. Unlike the IDs themselves, embeddings contain numerical features that later model layers can use.

That gives us the main path for this Level:

text
→ tokens
→ token IDs
→ model-ready sequence
→ embeddings + position information

A few words you will meet​

TermPlain meaning in this Level
tokenone piece of text handled as a unit by a tokenizer
vocabularythe saved list or mapping of token pieces and their IDs
token IDthe integer used to look up a token or its embedding row
unknown token (<unk>)a fallback used when a tokenizer cannot represent a piece directly
paddingextra placeholder positions added so sequences can share a length in a batch
maska set of yes/no-like markers telling later computation which positions should count
embeddinga learned numerical vector selected for a token ID
context windowthe maximum sequence length the model can use at one time

You will also compare characters, bytes, whole words, and subwords as possible text pieces. A subword is simply a reusable piece that can be smaller than a whole word, such as play + ing.

Learning goal​

By the end of this Level, you should be able to take a short string, trace every representation step, and explain where information may be preserved or lost.

You should be able to:

  • explain why characters, words, bytes, and subwords create different tradeoffs between vocabulary size and sequence length;
  • show how a vocabulary maps pieces to IDs and what information can be lost with <unk>;
  • explain why some special token IDs must stay consistent between the tokenizer, saved data, and model;
  • explain what padding, truncation, and masks do to a batch;
  • distinguish token IDs from embeddings;
  • explain why position information is needed in addition to token identity;
  • explain how tokenization affects how much text fits inside a fixed context window;
  • evaluate a tokenizer using several kinds of text instead of one convenient example.

When this Level uses the word contract, it means a set of choices that must agree with one another. For example, if the tokenizer says ID 3 means <eos> but the model assumes ID 3 means something else, the two parts do not share the same contract and the system is broken.

Learning path​

L4.1–L4.3 establish the basic problem: turn text into pieces, compare possible piece sizes, and make the vocabulary/unknown-token problem visible.

L4.4 — Build a Simple Tokenizer assembles those ideas into one small complete tokenizer.

L4.5–L4.7 add byte-level representation, learned subword pieces, token IDs, and special tokens. After L4.7 — Token IDs and Special Tokens, complete the Level 4 mini checkpoint.

L4.8–L4.12 connect tokenization to model input: padding and masks, embeddings, positions, and context windows.

L4.13–L4.14 evaluate the tokenizer on different kinds of text and turn the earlier pieces into an engineering review.

How to debug this Level​

When an encoded result looks strange, do not jump straight to the final model output. Keep the original text and trace the earliest place where the result becomes surprising:

source text
→ pieces
→ IDs
→ masks / positions
→ embeddings

For example, if one word becomes <unk>, the problem began before the embedding lookup. If the pieces are correct but the IDs are wrong, inspect the vocabulary mapping next.

Level Project​

After L4.14 — Tokenizer Evaluation Workshop, complete Tokenizer Workbench.

You will compare tokenization choices on fixed examples, verify that special tokens and masks stay consistent, inspect embedding/position alignment, and document one reproducible failure and fix.

Lesson actions

Completion is stored locally on this device.

View progress