Skip to main content

Level 5 project

Build a Mini Transformer

Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.

Launch lesson: Build a Mini Transformer

Prerequisite project: Tokenizer Workbench

Goal

Build a deterministic decoder-style mini Transformer from understandable pieces and prove the invariants that make it causal and composable.

Task

Complete the TODOs in projects/starters/l05/mini_transformer.py.

Your mini Transformer must:

  1. implement numerically stable softmax;
  2. normalize each token's feature vector;
  3. implement scaled causal multi-head self-attention with future probability exactly zero;
  4. reject an embedding width that is not divisible by the head count;
  5. preserve residual-stream shape through attention, feed-forward, and stacked blocks;
  6. assemble token + position embeddings, stacked blocks, final normalization, and vocabulary logits;
  7. keep earlier logits unchanged when only future token IDs change;
  8. expose attention weights so causal behavior can be inspected.

Validation

Run these commands from the downloaded Project folder or the public materials repository root.

python3 projects/tests/l05/validate_submission.py projects/starters/l05/mini_transformer.py

Rubric

Project ID: p05-mini-transformer Launch: after l05-16-build-a-mini-transformer Canonical dependency: p04-tokenizer-workbench

Score each criterion 0–4. A strong submission earns at least 16/20 and has no zero in Causality or Reproducibility.

CriterionLevel 5 exit-skill mapping4 — Strong evidence2 — Partial evidence0 — Missing/incorrect
Scaled causal multi-head attentionImplement scaled causal MHAScores are scaled, heads split/rejoin correctly, rows normalize, future mass is zero, divisibility checkedAttention mostly works but one invariant is weak or untestedFuture leakage or no multi-head implementation
Residual/norm/MLP blockExplain and implement residual/norm/MLP rolesPre-norm residual paths preserve shape and MLP is position-wise with explicit checksComponents exist but ordering/role explanation is incompleteBlock contract is absent or shape-invalid
Stacking and full architectureAssemble and stack Transformer blocksEmbeddings → stack → final norm → vocabulary logits are traced and depth preserves (T,C)Complete path exists but shape reasoning is weakNo complete mini Transformer
Debugging and evaluationDebug shape/mask errorsIntentional divisibility and future-leakage failures are reproduced; evidence → hypothesis → focused fix is clearOne failure is shown without strong causal evidenceOnly happy-path output
Reproducibility and communicationReproduce and explainStandard validator passes from clean Python 3.11+, no network/secrets, architecture and causal reasoning are clearRuns with undocumented assumptions or weak explanationCannot reproduce or relies on unrecorded external state

Required evidence

  • objective validator output;
  • architecture shape note;
  • causal prefix-invariance evidence;
  • future-attention-mass evidence;
  • intentional failure/debug note;
  • short explanation of why attention mixes positions while the feed-forward network acts position-wise.

Human review should reward reasoning and observable evidence, not byte-identical prose.

← Back to all projects