Skip to main content
L5.15

Stacking Transformer Blocks

Goal

Trace a residual stream through several Transformer blocks and verify that shape and causal invariants survive depth.

One block performs one round of routing and feature transformation. A stack repeats that contract. The values should change with depth; the interface should not:

(B,T,C) → block 1 → (B,T,C)
→ block 2 → (B,T,C)
→ ...

Depth does not retokenize the sequence or multiply T. It repeatedly refines the same-shaped residual stream.

Depth changes the computation history, not the interface​

After block 1, each token representation can contain information routed from legal context. Block 2 receives those already contextualized representations and performs another round of routing and feature transformation. Therefore later blocks are not simply repeating the exact same calculation on the original embeddings.

A useful way to think about depth is iterative refinement:

embedding stream
→ context/update round 1
→ context/update round 2
→ context/update round 3
→ ...

The shape can remain (B,T,C) throughout while the meaning carried by the feature values becomes progressively more task-specific.

This also explains why one causal leak is serious. If block 2 allows an earlier position to read a future token, block 3 receives a contaminated representation even if block 3's own mask is correct. Causality is therefore a property that must hold through every layer, not an initialization step performed once at the bottom of the model.

Later blocks receive representations that already contain routed context​

Suppose block 1 lets position 5 gather useful evidence from positions 1 and 3.

Block 2 does not receive the original isolated token embedding at position 5. It receives a representation that already carries some of that earlier context.

Block 2 can then route and transform this richer state again.

Depth therefore creates repeated rounds of:

read current representation
→ route context
→ transform features
→ preserve/update residual stream

This helps explain why inspecting only the first attention layer rarely tells the whole story of a deep model.

It also creates a debugging consequence: once an earlier block produces corrupted values, every later block receives those corrupted states. Find the earliest layer where an invariant breaks.

Predict

If six valid blocks each preserve `(B,T,C)`, what shape should the stack return?

Localize failures by layer index​

Open the notebook and run the code cell. It builds model=Stack(depth=4), prints depth: 4, prints the shape after each block, and ends with PASS: l05-15 stacked block checks.

  1. Confirm that every after block ... line shows (2, 6, 12).
  2. Change model=Stack(depth=4) to model=Stack(depth=1) and rerun. Then try depth=2. Before each run, predict how many after block lines will appear and what shape each will have.
  3. Record your observation: adding blocks adds computation, but never changes the (B, T, C) interface. That is what lets blocks be stacked.
  4. Restore depth=4. As a thought experiment, suppose one block in the middle forgot its causal mask. The shapes would still print (2, 6, 12) everywhere, so shape checks alone cannot catch that bug—you need the future-attention check from L5.8 at every layer.

Loading lab…

A decoder stack must preserve causality at every attention layer. One leaking layer is enough to let future information contaminate later representations.

Quick Check

1. What changes across blocks?
2. How many decoder attention layers must respect the causal mask?
3. A deep stack fails. What evidence is most useful first?

0 of 3 questions answered.

Explain it back​

Explain why increasing depth increases computation but does not require changing the residual-stream shape contract.

Key Takeaways

  • Transformer depth is repeated block composition over one residual stream.
  • Shape and causality must hold at every layer boundary.
  • Record layer index and invariant evidence to debug deep stacks.

Next Lesson

Next, connect token/position embeddings, the block stack, final normalization, and vocabulary projection into a complete mini Transformer architecture.

References

Lesson actions

Completion is stored locally on this device.

View progress