Skip to main content
L5.14

Transformer Block

Goal

Compose causal attention, normalization, residual connections, and a feed-forward network into one transformer block, then trace the shape through the block.

Two transformations, one continuing stream​

A language model does not use attention once and stop. It repeats a block that lets token representations:

  1. mix information across positions with causal attention;
  2. transform features at each position with a feed-forward network.

Residual connections keep a direct path for the existing representation while each sublayer contributes an update.

A pre-norm block can be pictured like this:

x
│
├── LayerNorm ── Causal Attention ──┐
│ │
└────────────────────────────────── + ── x1
│
├── LayerNorm ── Feed Forward ──┐
│ │
└────────────────────────────── + ── output

The + operations are residual additions.

Predict

If a transformer block receives shape (B, T, C), what shape should it return?

Read the forward pass as two updates​

The whole block can look surprisingly short:

def forward(self, x):
x = x + self.attn(self.ln1(x))
x = x + self.ff(self.ln2(x))
return x

Read the first line as:

normalize the current stream, compute an attention update, then add that update back.

Read the second line the same way for the feed-forward branch.

The feed-forward network temporarily expands the channel dimension, applies a nonlinearity, then returns to n_embd:

self.ff = nn.Sequential(
nn.Linear(n_embd, 4 * n_embd),
nn.GELU(),
nn.Linear(4 * n_embd, n_embd),
)

Returning to n_embd is not cosmetic. The residual addition requires the branch output to have the same (B, T, C) shape as the stream it is added to.

Check the block's shape rule​

  1. Run the Lab unchanged.
  2. Find the input and output shapes for the transformer block.
  3. Confirm that both are (B, T, C) for the same batch, time, and channel sizes.
  4. In the notebook's controlled experiment, make the final feed-forward projection return the wrong channel width.
  5. Observe that the residual addition fails because the shapes no longer match.
  6. Restore the original projection afterward.

Loading lab…

When a residual addition fails, inspect the two branch shapes immediately before +. Optimizer tuning cannot repair a shape mismatch.

Attention and feed-forward layers do different jobs​

Attention mixes information between token positions.

The feed-forward network applies the same learned channel transformation independently at each position. It changes the features of a token representation without directly combining different positions.

This distinction will matter in Level 6 when several blocks are stacked: each block alternates cross-position mixing with position-wise feature transformation while preserving one residual-stream shape.

Under the Hood

This curriculum uses pre-norm blocks: normalization happens before attention and before the feed-forward branch. Other transformer families place normalization differently. The important skill here is tracing exactly which tensor each operation receives and returns.

Quick Check

1. Why must the attention and feed-forward outputs return to n_embd channels?
2. Does the feed-forward network mix different token positions directly?
3. What is a useful way to read a residual sublayer?

0 of 3 questions answered.

Key Takeaways

  • A transformer block alternates attention and a position-wise feed-forward transformation.
  • Residual connections add learned updates to the continuing representation stream.
  • Each residual branch must return to (B, T, C).
  • Pre-norm means normalization happens before each sublayer in this design.
  • Debug residual failures by checking branch shapes before training behavior.

Next Lesson

Next, L5.15 — Stacking Transformer Blocks repeats this block and checks that the same stream shape and no-future rule survive depth.

References

Lesson actions

Completion is stored locally on this device.

View progress