Dot Products as Similarity
Goal
Compute small dot products by hand and explain how aligned signs, directions, and magnitudes influence attention scores.
Turn a query-key match into one score
From L5.3, remember the job: a query asks what is useful now, and each key is a candidate match. Attention needs one number for each query-key pair so it can compare those candidates before mixing values.
The dot product is that scoring rule. For two vectors, pair the components, multiply, then add:
[1, 2] · [3, 4] = 1×3 + 2×4 = 11
Against query [1, 0], key [1, 0] scores higher than [0, 1] because the active components align. Against [1, 1], key [1, -1] scores 0 because one positive and one negative contribution cancel.
Open one query-key score
Choose a candidate and inspect the component products that add up to its raw compatibility score.
[1, 0] · [1, 0] = 1×1 + 0×0 = 1.000
| Candidate | Key | Raw dot-product score |
|---|---|---|
| Aligned key | [1, 0] | 1.000 |
| Different-axis key | [0, 1] | 0.000 |
A dot product is a score, not a universal meaning detector
Consider query [2, 1].
key A = [2, 1] → 2×2 + 1×1 = 5
key B = [1, 2] → 2×1 + 1×2 = 4
key C = [-2,-1] → 2×(-2) + 1×(-1) = -5
The arithmetic tells you that A aligns most strongly with this query under the current learned coordinate system. It does not by itself tell you that A is “more similar in human meaning.” The vectors only become useful because training learns dimensions and projections that make these scores helpful for the task.
Magnitude matters too. [20,10] has the same direction as [2,1] but produces a much larger dot product. So there are two distinct ideas:
- alignment of components affects sign and relative score;
- vector magnitude affects score scale.
The next lesson's scaling factor addresses the second issue at the attention-logit level. When debugging, calculate a few component products by hand before blaming softmax; a wrong sign or transpose can be found much earlier.
Raw scores are relative, not calibrated meanings
A raw attention score such as 5 is not universally “good” or “similar.” The raw score comes from that query-key pair and its learned projections. Its effect on the final attention weights depends on how it compares with the other scores competing for the same query.
That gives a useful debugging boundary: compare component products and relative scores inside the same attention calculation, but do not treat raw dot products from different heads, layers, or inputs as if they shared one calibrated semantic scale. The next lesson will show how scaling changes these logits before softmax.
Predict
Make the score auditable
- Click Run. With
query = [1.0, 1.0], the scores arealigned 2.0,orthogonal-ish 0.0, andopposite -2.0. - Flip one component sign. Change
"aligned": [1.0, 1.0],to"aligned": [1.0, -1.0],. - Before running, predict the new
alignedscore. Work it out as two products: 1 × 1 and 1 × (-1). - Click Run. The score becomes
0.0: the two products,1and-1, cancel. If a result surprises you, inspect each product separately rather than jumping to softmax. - Now try magnitude instead of direction. Change the same key to
[2.0, 2.0]. It points the same way as the query, just twice as long. Run again: the score doubles to4.0. - Press Reset afterward.
Loading lab…
Quick Check
Transfer the idea
Explain why “higher dot product means more semantically similar” is too strong without specifying how the vectors were learned and scaled.
Key Takeaways
- Dot products turn pairwise vector alignment into scalar scores.
- Signs and magnitudes both affect raw scores.
- Component-by-component arithmetic is the smallest useful debug unit.
- Large score magnitudes motivate scaling before softmax.
Next Lesson
Next, scale the score matrix so increasing head dimension does not make softmax needlessly sharp.
References
- Vaswani et al., Attention Is All You Need.
- Luong, Pham, and Manning, Effective Approaches to Attention-based Neural Machine Translation.
Completion is stored locally on this device.