Positional Encodings
Goal
Compute and inspect a small positional encoding, track its [sequence_length, embedding_width] shape, and explain why its indexing must align with token embeddings.
A sequence position must become numerical information before the network can combine it with token features. A classic sinusoidal scheme uses sine and cosine waves at different rates.
The formula matters less at first than the shape and alignment:
token embeddings: [sequence_length, embedding_width]
position encodings: [sequence_length, embedding_width]
combined representation: same shape
If embedding width is 4, a position vector directly added to it must also have width 4.
Turn position numbers into vectors
A model width of C=8 means token embeddings have eight features. Position information therefore also needs a representation compatible with those eight features.
One approach is a learned position table:
position 0 → P0 = [ ... 8 values ... ]
position 1 → P1 = [ ... 8 values ... ]
position 2 → P2 = [ ... 8 values ... ]
Then:
input at position 2 = token_embedding + P2
Another approach computes position vectors from a formula, such as sinusoidal encodings.
The important invariant is the same: every position receives a distinguishable position signal with the same feature width as the token representation.
Learned and fixed encodings make different trade-offs
A learned table can adapt its position vectors during training, but it only has entries for the positions represented in the table.
A fixed mathematical encoding does not need one learned row per position and can define vectors beyond the training range, although that does not guarantee the model will use much longer contexts well.
Do not collapse these ideas into “fixed is always better for long context.” A model trained only on short sequences may still fail when asked to generalize far beyond them.
Position information must survive the whole input pipeline
Padding, special-token insertion, slicing, and batching can all shift or mask positions.
A correct positional-encoding formula does not protect you from giving it the wrong position index.
When debugging, trace a small sequence all the way from text to token IDs to positions to combined input vectors.
Predict
Inspect positions, not just formulas
The Lab builds a position table with positional_encoding(length=4, width=4) and prints one row per position.
- Predict
sin(0)andcos(0)before running. (They are0and1.) - Click Run. Row
0is[0.0, 1.0, 0.0, 1.0]: every sine column starts at0and every cosine column at1. - Compare rows
1,2, and3. The first two columns change quickly from row to row, while the last two change slowly (0.01,0.02,0.03). Different columns “tick” at different speeds, so every position gets a different pattern. - Change only the length:
positional_encoding(length=6, width=4). Before running, predict whether rows0–3will change. - Click Run. Rows
0–3are identical to before, and two new rows (4and5) appear. Each position's code depends only on the position and the width, not on how long the sequence is. - Press Reset afterward.
Loading lab…
Now deliberately create a width mismatch. A visible shape error is easier to debug than a silent indexing shift, so always write semantic dimension names beside numerical shapes.
Quick Check
Explain it back
Explain how a deterministic position vector can add useful information even though it was not learned from text.
Key Takeaways
- Position must be represented numerically for a sequence model.
- Position tables align by sequence position and embedding width.
- Matching shapes are necessary; matching token-position indexing is equally important.
- Sinusoidal encoding is one mechanism, not the definition of position itself.
Next Lesson
Next, connect representation length to a model's finite context capacity.
References
- Vaswani et al., Attention Is All You Need.
- Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding.
Completion is stored locally on this device.