Skip to main content
L4.2

Characters, Words, and Subwords

Goal

By the end of this lesson, you can compare three ways to split text into token pieces and explain why smaller pieces usually need more sequence positions while larger pieces need a larger vocabulary to cover many words.

One word can be split several ways​

Take the word playing. A tokenizer could represent the same text using character-sized pieces, one whole-word piece, or reusable subword pieces.

A position here means one token slot in the sequence sent toward the model. Compare how many positions each choice needs:

Tokenization: one text, different piece sizes

Compare character, whole-word, and subword choices for the same text. Smaller pieces are easier to reuse, but they usually take more sequence positions.

Same source textplaying
StrategyPiecesPositions
Charactersp | l | a | y | i | n | g7
Whole wordplaying1
Subwordsplay | ing2

A whole-word tokenizer uses a complete word as one token when that word exists in its vocabulary. This is compact for familiar words, but a vocabulary cannot realistically contain every name, spelling variation, or newly invented word.

A character tokenizer uses much smaller pieces. It can build unfamiliar spellings from a small reusable set of characters, but common words take many more positions.

A subword tokenizer uses reusable pieces that are often larger than one character but smaller than a complete word. For example, playing can reuse play and ing. The split does not always match the word parts a grammar teacher would choose; the goal is useful reuse for the tokenizer.

So there is a tradeoff:

larger pieces
→ fewer positions for familiar text
→ more distinct pieces may be needed in the vocabulary

smaller pieces
→ more reusable pieces
→ more positions are often needed for the same text

Here coverage means whether the tokenizer has a way to represent the text it receives. A strategy with good coverage can still encode unfamiliar text instead of failing only because the complete word was unseen.

Predict

A spelling game contains many invented words. Which piece size is least likely to fail only because a complete word is new?

Compare the same text in the Lab​

The Lab prints character, word, and simple subword versions of the same text.

  1. Click Run without editing anything.
  2. For players replayed, compare the three lines. Pay attention to both count= and pieces=.
  3. Notice that the simple subword rule splits the suffixes, producing pieces such as play + ers and replay + ed.
  4. Find:
text = "players replayed"
  1. Change only the text to:
text = "play replay"
  1. Before running, predict the direction of the change. The word representation still needs 2 positions. The simple subword representation should also need only 2 positions now because neither word has one of the suffixes that the starter splits off.
  2. Click Run and compare all three count= values with the first run.
  3. Restore text = "players replayed".

Loading lab…

This small program is only one toy subword rule. Real subword tokenizers learn or choose reusable pieces from much larger text collections. The point of the Lab is the tradeoff between piece size, reuse, and sequence length—not to claim that this suffix rule is a production tokenizer.

Quick Check

1. Why are character-token sequences often longer?
2. What is a weakness of relying only on a fixed whole-word vocabulary?
3. Why use subword pieces?

0 of 3 questions answered.

Transfer the idea​

Suppose a chat product works well on common English words. Why would that alone not prove that its tokenizer works efficiently for names, emoji, compounds, misspellings, or other writing systems?

The answer should mention the kind of text tested. A tokenizer's behavior depends on the inputs it must actually represent.

Key Takeaways

  • The same text can be split into character, whole-word, or subword tokens.
  • Smaller pieces are broadly reusable but usually create longer sequences.
  • Larger pieces can shorten familiar text but require broader vocabulary coverage.
  • Subwords aim for a useful middle ground between reuse and sequence length.
  • Judge tokenization choices on the kinds of text the system actually needs to handle.

Next Lesson

Next, in L4.3 — Vocabulary and Unknown Tokens, you will make the piece-to-ID mapping explicit and see what happens when a vocabulary cannot represent a piece directly.

References

Lesson actions

Completion is stored locally on this device.

View progress