Skip to main content
L5.2

Words Looking at Words

Goal

Explain attention as a weighted information lookup between token positions and read one normalized attention row without needing to implement the full attention mechanism yet.

A token often needs context from other positions. In the sentence:

the animal did not cross the road because it was tired

it is easier to interpret if information from animal receives more weight than unrelated words.

Attention gives each query position a mixture of information from legal context positions. The important new idea in this lesson is the mixture itself:

output = weight1 × value1 + weight2 × value2 + ...

The weights are non-negative and, after normalization, add up to 1 across the positions that are allowed to contribute.

Read one attention row​

A tiny decoder attention map

Select a row and read it as a distribution over context positions. Future cells are already hidden here; L5.8 — Causal Masking will explain how that no-future rule is implemented.

Query: predicts · strongest visible link: model
query ↓ / key →
the
model
predicts
next
the
100%
model
35%
65%
predicts
15%
45%
40%
next
10%
30%
35%
25%

For the row belonging to predicts, the visible weights are 0.15, 0.45, and 0.40. They add to 1.0.

That means the output at this position is built from a weighted mixture rather than one fixed average of all context.

Predict

After normalization, what should the legal attention weights in one row add up to?

Use the Lab only to inspect weighting behavior​

The browser Lab contains scores = [0.2, 1.0, 0.4, 2.0] and prints one normalized row for each query position. The starter already handles the decoder's future restriction. You do not need to understand the masking code yet; L5.8 — Causal Masking will build and test that no-future rule directly.

  1. Click Run unchanged.
  2. Find the line beginning with 2. That row corresponds to query index 2, so only score positions 0, 1, and 2 are legal and the fourth weight is 0.0.
  3. Confirm the printed sum= is 1.0.
  4. In scores, change only the first value from 0.2 to 2.0:
scores = [2.0, 1.0, 0.4, 2.0]
  1. Before running, predict that query row 2 should give more weight to position 0, because that legal score became larger.
  2. Click Run and compare only row 2 with the first run. Confirm the legal weights still sum to 1.0 and the fourth future position stays at 0.0.
  3. Restore the first score to 0.2.

Loading lab…

This is enough for now: scores change → normalized weights change → the weighted information mixture changes.

What you are not expected to know yet​

A full Transformer computes those scores from learned vector roles called queries, keys, and values. It also scales scores, applies a causal mask for decoder models, and may split feature channels across several heads.

Those are not prerequisites for this lesson. The next lessons introduce them one at a time:

  • L5.3 — Queries, Keys, and Values
  • L5.4 — Dot Products as Similarity
  • L5.5 — Scaled Dot-Product Attention
  • L5.6 — Attention Weights and Weighted Values
  • L5.8 — Causal Masking
  • L5.9 — Multi-Head Attention

Seeing the names now is a map, not an instruction to understand the implementation early.

Attention weights are evidence, not a complete explanation​

A weight row is useful because it shows where one attention operation routed information. It does not prove that a token or head contains a complete human-readable reason for the model's final output. Later layers, value directions, residual connections, and nonlinear transformations also matter.

Quick Check

1. What does one normalized attention row represent?
2. A legal score becomes much larger while the other legal scores stay fixed. What should usually happen after softmax?

0 of 2 questions answered.

Key Takeaways

  • Attention builds a weighted mixture of information from context positions.
  • A normalized legal weight row sums to 1.
  • Changing one score changes the distribution of weight and therefore the mixture.
  • The Lab previews a decoder restriction, but masking and multi-head implementation are taught later rather than assumed here.
  • Attention weights are useful evidence, not a complete account of model behavior.

Next Lesson

Next, you will make the scoring rule concrete with L5.3 — Queries, Keys, and Values, separating what a position asks for, what each position advertises for matching, and what information is actually carried forward.

References

Lesson actions

Completion is stored locally on this device.

View progress