Embeddings for Tokens
Goal
Use token IDs to select embedding rows, track the resulting sequence shape, and explain why learned vector values—not ID magnitude—carry model features.
Continue with the retained IDs from the previous lesson. If ID 3 selects [0.2, -0.4, 0.7], the network receives that vector. The integer 3 is only the row address.
For a sequence of 5 IDs and embedding width 3, lookup produces shape (5, 3): one vector per token position.
This separation is important:
tokenizer: text piece → ID
embedding table: ID → learned vector
Changing the tokenizer mapping while reusing the old table can connect the wrong word to a perfectly valid vector.
An embedding is a learned table lookup
Suppose the vocabulary has 1,000 token IDs and the model width is 8.
The embedding table has shape:
(1000, 8)
Token ID 17 retrieves row 17:
17 → [0.12, -0.44, 0.03, ... 8 values total]
For a sequence of four token IDs, the result has shape:
(4,) IDs → (4,8) vectors
For a batch of three such sequences:
(B,T) = (3,4) → (B,T,C) = (3,4,8)
This is the first point where discrete token identities become continuous vectors that later neural-network layers can transform.
The coordinates are learned, not assigned by vocabulary order
Token ID 18 is not automatically “close in meaning” to token ID 17. Integer IDs are just lookup addresses.
Meaningful geometric relationships emerge only if training updates embedding rows in ways that support the prediction objective.
Two tokens used in similar contexts may eventually receive useful related representations, but that is a learned result, not a property of their IDs.
One token can develop different contextual meaning later
At the embedding lookup stage, token ID 17 always retrieves the same base row for a fixed checkpoint.
Later Transformer layers can turn that same starting vector into different contextual representations depending on surrounding tokens.
So distinguish:
- token embedding: the learned starting vector for an ID;
- contextual hidden state: the later vector after attention and other transformations.
This distinction becomes important when people casually say “the embedding for this word.” Ask which stage they mean.
Predict
Trace lookup instead of guessing
The Lab's EMBEDDINGS table has 4 rows (IDs 0–3) and width 3.
- Before running, use the table to predict the vectors for
token_ids = [1, 3, 2]: copy row 1, then row 3, then row 2. - Click Run and compare. The output should be
[[0.2, -0.4, 0.7], [0.5, 0.1, -0.2], [-0.1, 0.9, 0.3]]: three IDs in, three vectors of width 3 out, so the shape is(3, 3). - Change only the last ID:
token_ids = [1, 3, 3]. Before running, predict the last two vectors. - Click Run. The last two vectors are identical,
[0.5, 0.1, -0.2]. The same ID always selects the same starting row. - Now try an ID the table does not have:
token_ids = [1, 3, 4]. Run it. The Lab stops withIndexError: token id is outside the embedding table. This is what happens when a tokenizer produces more IDs than the model's embedding table was built for. - Press Reset afterward.
Loading lab…
Inspect both the lookup result and its shape. If a token appears to have the “wrong embedding,” check the vocabulary-to-ID mapping before changing the embedding values.
Quick Check
Explain it back
Explain why an embedding is a learned representation but a token ID is not. Include the lookup table in your explanation.
Key Takeaways
- IDs select rows; embedding vectors carry learnable features.
- Sequence length and embedding width become separate dimensions.
- Numeric ID distance is not semantic distance.
- Tokenizer versions and embedding-row meanings form one compatibility contract.
Next Lesson
Next, add a second kind of information: where each token occurs in the sequence.
References
- Kudo and Richardson, SentencePiece.
- Bengio et al., A Neural Probabilistic Language Model.
Completion is stored locally on this device.