Feed-Forward Networks
Goal
Trace the position-wise feed-forward network through expansion, nonlinearity, and projection back to the residual-stream width.
Attention lets one position gather information from other positions. The feed-forward network does a different job: it transforms the features inside each position independently, using the same learned rule at every time step.
A common pattern is:
C → larger hidden width → nonlinearity → C
Returning to C = n_embd is not cosmetic; the output must be addable back to the residual stream.
Follow one position while the others wait
Take a residual stream with three token positions:
token 0 → [ ... C features ... ]
token 1 → [ ... C features ... ]
token 2 → [ ... C features ... ]
The feed-forward network applies the same parameter matrices to each row independently. Token 0's output does not read token 1's features inside this sublayer. Cross-position information may already be present in token 0 because attention mixed it earlier, but the MLP itself transforms one position at a time.
This distinction helps you reason about failures. If swapping the order of two positions changes an MLP output even though their input feature vectors are unchanged, some unintended cross-position operation may have entered the branch.
Why expand before projecting back? A simple C → C linear map can only perform one linear transformation. A wider hidden layer plus a nonlinearity can build combinations that cannot be collapsed into one linear matrix. The final → C projection then converts that richer feature computation back into the common residual-stream interface.
Keep the two branch roles separate
Use the contrast as a debugging rule: attention mixes positions; the MLP transforms features. Attention can change one token because information arrived from another position. The MLP can change that token's feature pattern without directly reading another position.
For C = 8, one MLP might use:
8 → 32 → 8
The same learned matrices are applied independently at every token position. If the wrong token is being consulted, inspect attention; if the right contextual information is present but its feature transformation looks wrong, inspect the MLP branch.
Predict
Separate position mixing from feature transformation
The Lab's feed_forward follows the C → 2C → C pattern with tiny fixed teaching weights: it expands each feature x into two hidden units, x and -x, applies ReLU, and then combines each pair with weights 1.0 and -0.5.
- Click Run. Read
input width: 3 | hidden width: 6 | output width: 3. The branch expands, then returns to the residual width. - Compare inputs and outputs:
[1.0, -2.0, 0.5] -> [1.0, -1.0, 0.5]. Positive features pass through unchanged, while negative features are halved. The same input value is treated differently depending on its sign—that is the nonlinearity at work. - Swap the two samples:
samples = [[-1.0, 3.0, -4.0], [1.0, -2.0, 0.5]]. Run again. Each position gets exactly the same output as before, just in a different order. The feed-forward network never looks at other positions. - Press Reset.
Loading lab…
- Now remove the nonlinearity. Change
hidden = [relu(h) for h in hidden]tohidden = hidden. - Before running, predict what happens to positive and negative features.
- Click Run. Every feature is now multiplied by the same number,
1.5:[1.0, -2.0, 0.5] -> [1.5, -3.0, 0.75]. Without ReLU, the expand and contract steps collapse into one simple linear scaling. Two linear maps composed together are still one linear map; the nonlinear middle stage is what gives the branch richer feature transformation capacity. - Press Reset afterward.
Quick Check
Explain it back
Use the phrase “attention mixes positions; MLP transforms features” and explain what that distinction means in (B,T,C) terms.
Key Takeaways
- Attention and the feed-forward network have different jobs.
- The MLP is shared across positions but does not directly mix them.
- Nonlinearity makes the two-layer transformation more expressive.
- The branch projects back to the residual-stream width.
Next Lesson
Next, assemble normalization, causal attention, residual paths, and the feed-forward network into one traceable pre-norm block.
References
- Vaswani et al., Attention Is All You Need.
Completion is stored locally on this device.