Skip to main content
L5.12

Feed-Forward Networks

Goal

Trace the position-wise feed-forward network through expansion, nonlinearity, and projection back to the residual-stream width.

Attention lets one position gather information from other positions. The feed-forward network does a different job: it transforms the features inside each position independently, using the same learned rule at every time step.

A common pattern is:

C → larger hidden width → nonlinearity → C

Returning to C = n_embd is not cosmetic; the output must be addable back to the residual stream.

Follow one position while the others wait​

Take a residual stream with three token positions:

token 0 → [ ... C features ... ]
token 1 → [ ... C features ... ]
token 2 → [ ... C features ... ]

The feed-forward network applies the same parameter matrices to each row independently. Token 0's output does not read token 1's features inside this sublayer. Cross-position information may already be present in token 0 because attention mixed it earlier, but the MLP itself transforms one position at a time.

This distinction helps you reason about failures. If swapping the order of two positions changes an MLP output even though their input feature vectors are unchanged, some unintended cross-position operation may have entered the branch.

Why expand before projecting back? A simple C → C linear map can only perform one linear transformation. A wider hidden layer plus a nonlinearity can build combinations that cannot be collapsed into one linear matrix. The final → C projection then converts that richer feature computation back into the common residual-stream interface.

Keep the two branch roles separate​

Use the contrast as a debugging rule: attention mixes positions; the MLP transforms features. Attention can change one token because information arrived from another position. The MLP can change that token's feature pattern without directly reading another position.

For C = 8, one MLP might use:

8 → 32 → 8

The same learned matrices are applied independently at every token position. If the wrong token is being consulted, inspect attention; if the right contextual information is present but its feature transformation looks wrong, inspect the MLP branch.

Predict

Why does the final feed-forward projection usually return to `n_embd`?

Separate position mixing from feature transformation​

The Lab's feed_forward follows the C → 2C → C pattern with tiny fixed teaching weights: it expands each feature x into two hidden units, x and -x, applies ReLU, and then combines each pair with weights 1.0 and -0.5.

  1. Click Run. Read input width: 3 | hidden width: 6 | output width: 3. The branch expands, then returns to the residual width.
  2. Compare inputs and outputs: [1.0, -2.0, 0.5] -> [1.0, -1.0, 0.5]. Positive features pass through unchanged, while negative features are halved. The same input value is treated differently depending on its sign—that is the nonlinearity at work.
  3. Swap the two samples: samples = [[-1.0, 3.0, -4.0], [1.0, -2.0, 0.5]]. Run again. Each position gets exactly the same output as before, just in a different order. The feed-forward network never looks at other positions.
  4. Press Reset.

Loading lab…

  1. Now remove the nonlinearity. Change hidden = [relu(h) for h in hidden] to hidden = hidden.
  2. Before running, predict what happens to positive and negative features.
  3. Click Run. Every feature is now multiplied by the same number, 1.5: [1.0, -2.0, 0.5] -> [1.5, -3.0, 0.75]. Without ReLU, the expand and contract steps collapse into one simple linear scaling. Two linear maps composed together are still one linear map; the nonlinear middle stage is what gives the branch richer feature transformation capacity.
  4. Press Reset afterward.

Quick Check

1. Which sublayer directly mixes different time positions?
2. Why expand to a wider hidden layer?
3. What output width must the residual branch return?

0 of 3 questions answered.

Explain it back​

Use the phrase “attention mixes positions; MLP transforms features” and explain what that distinction means in (B,T,C) terms.

Key Takeaways

  • Attention and the feed-forward network have different jobs.
  • The MLP is shared across positions but does not directly mix them.
  • Nonlinearity makes the two-layer transformation more expressive.
  • The branch projects back to the residual-stream width.

Next Lesson

Next, assemble normalization, causal attention, residual paths, and the feed-forward network into one traceable pre-norm block.

References

Lesson actions

Completion is stored locally on this device.

View progress