본문으로 건너뛰기
L5.1

Why Attention

Goal

By the end of this lesson, you can explain why a token may need different amounts of information from different earlier positions, and how attention represents that choice with weights.

Start with one sentence​

Consider:

the cat chased the toy because it rolled

Suppose we are building a new representation for the token it.

The earlier words do not all seem equally useful. In this sentence, toy may be more relevant than the or because for understanding what it refers to.

One simple but weak strategy would be to average information from every earlier position equally. If three pieces received equal weights, we might have:

weight on piece A = 1/3
weight on piece B = 1/3
weight on piece C = 1/3

A weight is just a number saying how much influence one piece has in the mixture.

Attention improves on the fixed average by calculating different weights from the current representations. A more relevant position can receive a larger weight.

That is what attention means at this stage of the course:

build a new representation by mixing information from other allowed positions, using larger or smaller weights depending on the current content.

We say allowed positions because some tasks place rules on which positions may be used. In the next-token models you will build later, a position may use itself and earlier tokens but not future tokens. Level 5 will introduce that causal rule explicitly in L5.8 — Causal Masking.

What gets mixed?​

Use three simple numbers as stand-ins for information stored at three positions:

values = [1, 5, 9]
weights = [0.2, 0.3, 0.5]

The weighted mixture is:

1×0.2 + 5×0.3 + 9×0.5 = 6.2

The numbers being mixed are called values in attention. Later, a value will be a vector rather than one number, but the idea is the same.

The weights in a valid attention row are non-negative and normally add up to 1, so the output is a mixture of the available values.

From a fixed summary to a different summary for each position​

The important contrast is not “attention versus no context.” Earlier sequence models can carry context too. The difference is that attention can build a different context mixture for each query position.

Suppose the earlier value vectors are simplified to three numbers:

cat → 8
toy → 2
rolled → 6

A fixed average always gives (8 + 2 + 6) / 3 = 5.33, no matter which current token is asking for context. But two query positions can need different evidence:

query A weights: [0.8, 0.1, 0.1] → 7.2
query B weights: [0.1, 0.7, 0.2] → 3.4

The values did not change. The routing weights changed because the query changed. That is the useful mental model to carry forward.

A common misconception is that a large attention weight means “this word caused the final answer.” It means something narrower: in this attention operation, more of that position's value representation is routed into the current mixture. Later layers and residual paths still transform the result.

Predict

Three positions always receive equal weights `[1/3, 1/3, 1/3]`. What can this fixed mixture not do?

Compare equal and focused weights in the Lab​

The Lab contains two score patterns. Softmax turns each score pattern into weights that add up to 1. You will study softmax more closely later; for now, read it as the step that converts comparison scores into usable mixture weights.

  1. Click Run without editing anything.
  2. Read uniform weights:. The starting scores [0.0, 0.0, 0.0] give all three positions equal weights.
  3. Read focused weights:. The scores [2.0, 0.0, 0.0] give the first position a larger weight.
  4. The values are [1.0, 5.0, 9.0]. Compare uniform output: with focused output: and notice that putting more weight on the first value pulls the output toward 1.0.
  5. Find:
focused = softmax([2.0, 0.0, 0.0])
  1. Change only 2.0 to 4.0.
  2. Before running, predict that the first focused weight will become even larger and focused output: will move even closer to 1.0.
  3. Click Run and compare the new focused weights and output with the first run.
  4. Restore 2.0 before moving on.

Loading lab…

The useful evidence is not just the final output. Inspect the weight row itself:

  • Are all weights non-negative?
  • Do they add up to about 1?
  • When one score becomes much larger, does its weight become larger too?

These checks make the mechanism visible before it is placed inside a larger Transformer.

Quick Check

1. Why can attention be more useful than always averaging every available position equally?
2. What do equal weights `[1/3, 1/3, 1/3]` do to three values?
3. If one comparison score becomes much larger while the others stay fixed, what direct evidence should you inspect?

0 of 3 questions answered.

Explain it back​

Explain attention without saying only “the model looks at words.” Your explanation should name:

  1. the information available at several positions;
  2. the weights that decide how much each position contributes;
  3. the weighted mixture produced for the current position.

Key Takeaways

  • Attention mixes information from several allowed positions.
  • A weight says how much influence one position has in that mixture.
  • Equal weights form an average and cannot express relevance differences.
  • Attention computes weights from the current representations rather than fixing them forever.
  • Inspect the weight row before blaming a much later model output.

Next Lesson

Next, in L5.2 — Words Looking at Words, you will make this weighted-mixture idea concrete with a small sequence and see how one position can assign different weights to other positions.

References

Lesson actions

Completion is stored locally on this device.

View progress