Skip to main content
L5.16

Build a Mini Transformer

Goal

Assemble the smallest complete decoder-style Transformer that maps token IDs to vocabulary logits while preserving explicit shape and causal invariants.

The whole Level now becomes one pipeline:

IDs (B,T)
→ token + position embeddings (B,T,C)
→ causal Transformer blocks (B,T,C)
→ final normalization (B,T,C)
→ vocabulary projection (B,T,V)

This architecture is not yet a trained language model. It is the computation that Level 6 will train for next-token prediction.

Read the model as three interfaces​

It helps to separate the full decoder into three questions.

1. Representation interface: How do integer IDs become position-aware feature vectors?
Token and position embeddings produce the residual stream (B,T,C).

2. Transformation interface: How are those feature vectors repeatedly updated without future leakage?
The causal Transformer stack preserves (B,T,C) while changing its values.

3. Prediction interface: How does one hidden vector become a next-token prediction?
The language-model head maps each C-wide vector to V raw vocabulary scores.

For a concrete example with B=2, T=4, C=8, and V=20:

IDs: (2,4)
hidden: (2,4,8)
logits: (2,4,20)

There are 2×4 separate prediction positions, and each receives 20 logits. Nothing in this shape says which token is correct yet; Level 6 will create shifted targets and a loss that supplies that training signal.

A model can pass every shape check and still be semantically wrong. Prefix invariance complements shape checks by testing the causal behavior that dimensions alone cannot reveal.

Keep an explicit shape ledger while assembling the model​

For a tiny configuration:

B = 2
T = 4
C = 8
V = 20
H = 2

write down the expected interfaces:

token ids: (2,4)
token embeddings: (2,4,8)
position signal: (2,4,8)
residual stream: (2,4,8)
each block output: (2,4,8)
final logits: (2,4,20)

This ledger catches an entire class of bugs before training.

But the ledger cannot check causality. A model can preserve every shape while still allowing an earlier token to use future information.

Finish with both interface evidence and behavioral evidence: expected tensor shapes, plus a prefix-invariance check showing that changing a future suffix does not alter earlier causal outputs.

Predict

Two input sequences share the same prefix but differ only in later tokens. What should a correct causal model do at the shared earlier positions?

Assemble and test contracts, not just modules​

Open the notebook and run it once unchanged. Find the line model=MiniTransformer(vocab_size=23, context=8, width=16, heads=4, depth=2). The run prints ids shape: (1, 5) | logits shape (B, T, V): (1, 5, 23) | blocks: 2 and ends with PASS: l05-16 mini transformer checks.

Then change one setting in that line at a time, predict which shapes change, and rerun:

  1. vocab_size=23 → vocab_size=30: only the last logits dimension, V, should change, to (1, 5, 30).
  2. depth=2 → depth=3: the logits shape stays (1, 5, 23) and blocks: becomes 3. More blocks add computation, not new output dimensions.
  3. heads=4 → heads=3: the notebook should stop with n_embd must be divisible by n_head, because 16 features cannot be split evenly into 3 heads.
  4. context=8 → context=4: the five-token input is now longer than the model's context, so the notebook stops with sequence exceeds context.

Restore every setting after each experiment, so only one thing changes at a time.

Loading lab…

Keep a short evidence checklist:

  • IDs lie inside vocabulary range;
  • embedding output is (B,T,C);
  • C is divisible by head count;
  • future attention mass is zero;
  • each block preserves (B,T,C);
  • logits are (B,T,V);
  • prefix outputs do not depend on later tokens.

Quick Check

1. What enters the first Transformer block?
2. What shape does the vocabulary head produce from `(B,T,C)`?
3. What does prefix invariance test?

0 of 3 questions answered.

Explain it back​

Explain the architecture from IDs to logits and name three independent invariants that would catch different implementation errors.

Key Takeaways

  • The mini Transformer connects Level 4 representations to a causal decoder computation.
  • Shape contracts make module composition auditable.
  • Causality must hold end to end, not only inside one attention row.
  • Level 6 will add targets, loss, optimization, checkpoints, validation, and generation.

Next Lesson

Complete Build a Mini Transformer. Then Level 6 will train this decoder-style architecture as a tiny language model.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Build a Mini Transformer

View progress