Stacking Transformer Blocks
Goal
Trace a residual stream through several Transformer blocks and verify that shape and causal invariants survive depth.
One block performs one round of routing and feature transformation. A stack repeats that contract. The values should change with depth; the interface should not:
(B,T,C) → block 1 → (B,T,C)
→ block 2 → (B,T,C)
→ ...
Depth does not retokenize the sequence or multiply T. It repeatedly refines the same-shaped residual stream.
Depth changes the computation history, not the interface
After block 1, each token representation can contain information routed from legal context. Block 2 receives those already contextualized representations and performs another round of routing and feature transformation. Therefore later blocks are not simply repeating the exact same calculation on the original embeddings.
A useful way to think about depth is iterative refinement:
embedding stream
→ context/update round 1
→ context/update round 2
→ context/update round 3
→ ...
The shape can remain (B,T,C) throughout while the meaning carried by the feature values becomes progressively more task-specific.
This also explains why one causal leak is serious. If block 2 allows an earlier position to read a future token, block 3 receives a contaminated representation even if block 3's own mask is correct. Causality is therefore a property that must hold through every layer, not an initialization step performed once at the bottom of the model.
Later blocks receive representations that already contain routed context
Suppose block 1 lets position 5 gather useful evidence from positions 1 and 3.
Block 2 does not receive the original isolated token embedding at position 5. It receives a representation that already carries some of that earlier context.
Block 2 can then route and transform this richer state again.
Depth therefore creates repeated rounds of:
read current representation
→ route context
→ transform features
→ preserve/update residual stream
This helps explain why inspecting only the first attention layer rarely tells the whole story of a deep model.
It also creates a debugging consequence: once an earlier block produces corrupted values, every later block receives those corrupted states. Find the earliest layer where an invariant breaks.
Predict
Localize failures by layer index
Open the notebook and run the code cell. It builds model=Stack(depth=4), prints depth: 4, prints the shape after each block, and ends with PASS: l05-15 stacked block checks.
- Confirm that every
after block ...line shows(2, 6, 12). - Change
model=Stack(depth=4)tomodel=Stack(depth=1)and rerun. Then trydepth=2. Before each run, predict how manyafter blocklines will appear and what shape each will have. - Record your observation: adding blocks adds computation, but never changes the
(B, T, C)interface. That is what lets blocks be stacked. - Restore
depth=4. As a thought experiment, suppose one block in the middle forgot its causal mask. The shapes would still print(2, 6, 12)everywhere, so shape checks alone cannot catch that bug—you need the future-attention check from L5.8 at every layer.
Loading lab…
A decoder stack must preserve causality at every attention layer. One leaking layer is enough to let future information contaminate later representations.
Quick Check
Explain it back
Explain why increasing depth increases computation but does not require changing the residual-stream shape contract.
Key Takeaways
- Transformer depth is repeated block composition over one residual stream.
- Shape and causality must hold at every layer boundary.
- Record layer index and invariant evidence to debug deep stacks.
Next Lesson
Next, connect token/position embeddings, the block stack, final normalization, and vocabulary projection into a complete mini Transformer architecture.
References
- Vaswani et al., Attention Is All You Need.
Completion is stored locally on this device.