본문으로 건너뛰기
L4.9

Embeddings for Tokens

Goal

Use token IDs to select embedding rows, track the resulting sequence shape, and explain why learned vector values—not ID magnitude—carry model features.

Continue with the retained IDs from the previous lesson. If ID 3 selects [0.2, -0.4, 0.7], the network receives that vector. The integer 3 is only the row address.

For a sequence of 5 IDs and embedding width 3, lookup produces shape (5, 3): one vector per token position.

This separation is important:

tokenizer: text piece → ID
embedding table: ID → learned vector

Changing the tokenizer mapping while reusing the old table can connect the wrong word to a perfectly valid vector.

An embedding is a learned table lookup​

Suppose the vocabulary has 1,000 token IDs and the model width is 8.

The embedding table has shape:

(1000, 8)

Token ID 17 retrieves row 17:

17 → [0.12, -0.44, 0.03, ... 8 values total]

For a sequence of four token IDs, the result has shape:

(4,) IDs → (4,8) vectors

For a batch of three such sequences:

(B,T) = (3,4) → (B,T,C) = (3,4,8)

This is the first point where discrete token identities become continuous vectors that later neural-network layers can transform.

The coordinates are learned, not assigned by vocabulary order​

Token ID 18 is not automatically “close in meaning” to token ID 17. Integer IDs are just lookup addresses.

Meaningful geometric relationships emerge only if training updates embedding rows in ways that support the prediction objective.

Two tokens used in similar contexts may eventually receive useful related representations, but that is a learned result, not a property of their IDs.

One token can develop different contextual meaning later​

At the embedding lookup stage, token ID 17 always retrieves the same base row for a fixed checkpoint.

Later Transformer layers can turn that same starting vector into different contextual representations depending on surrounding tokens.

So distinguish:

  • token embedding: the learned starting vector for an ID;
  • contextual hidden state: the later vector after attention and other transformations.

This distinction becomes important when people casually say “the embedding for this word.” Ask which stage they mean.

Predict

Token IDs 4 and 40 are numerically far apart. What does that tell you about their learned semantic similarity?

Trace lookup instead of guessing​

The Lab's EMBEDDINGS table has 4 rows (IDs 0–3) and width 3.

  1. Before running, use the table to predict the vectors for token_ids = [1, 3, 2]: copy row 1, then row 3, then row 2.
  2. Click Run and compare. The output should be [[0.2, -0.4, 0.7], [0.5, 0.1, -0.2], [-0.1, 0.9, 0.3]]: three IDs in, three vectors of width 3 out, so the shape is (3, 3).
  3. Change only the last ID: token_ids = [1, 3, 3]. Before running, predict the last two vectors.
  4. Click Run. The last two vectors are identical, [0.5, 0.1, -0.2]. The same ID always selects the same starting row.
  5. Now try an ID the table does not have: token_ids = [1, 3, 4]. Run it. The Lab stops with IndexError: token id is outside the embedding table. This is what happens when a tokenizer produces more IDs than the model's embedding table was built for.
  6. Press Reset afterward.

Loading lab…

Inspect both the lookup result and its shape. If a token appears to have the “wrong embedding,” check the vocabulary-to-ID mapping before changing the embedding values.

Quick Check

1. What does the embedding layer do with a token ID?
2. What shape results from 5 token positions and embedding width 3?
3. What must remain aligned if an embedding table is reused?

0 of 3 questions answered.

Explain it back​

Explain why an embedding is a learned representation but a token ID is not. Include the lookup table in your explanation.

Key Takeaways

  • IDs select rows; embedding vectors carry learnable features.
  • Sequence length and embedding width become separate dimensions.
  • Numeric ID distance is not semantic distance.
  • Tokenizer versions and embedding-row meanings form one compatibility contract.

Next Lesson

Next, add a second kind of information: where each token occurs in the sequence.

References

Lesson actions

Completion is stored locally on this device.

View progress