Mini Checkpoint — Through Sampling Text
Complete this checkpoint after L6.9 — Sampling Text. Keep the trained checkpoint fixed unless a question explicitly asks about training.
1. Shape trace
For B=2, T=16, C=48, and vocabulary V=64, write:
- token-ID shape;
- residual-stream shape after a Transformer block;
- logits shape;
- why the last dimension changes from
CtoVat the vocabulary head.
2. Next-token alignment
For [4,8,3,9,2] with block_size=4, write input and target rows. Then create one wrong alignment and explain which incorrect task it would train.
3. Data boundary
Explain why the source text should be divided into train/validation regions before overlapping windows are constructed. Give one example of a near-duplicate leakage path.
4. Reproducible checkpoint evidence
List at least eight pieces of state/metadata needed to understand and resume a run. Include model config, optimizer state, step, tokenizer/data identities, seed policy, and validation evidence.
5. Loss diagnosis
Training loss falls through step 100; validation loss reaches its minimum at step 60 and rises afterward. Which checkpoint would you inspect first under a minimum-validation-loss rule, and why?
6. Prefix to one sampled token
In 5–8 sentences, explain:
prefix IDs → embeddings/positions → causally masked blocks → logits → next-token distribution → sampled ID → extended prefix.
State clearly which part is model computation and which part is the sampling decision.
Check your reasoning after you try
- Shapes: token IDs are
(2, 16), the residual stream is(2, 16, 48), and logits are(2, 16, 64). The vocabulary head changes the last dimension because it scores every vocabulary token. - Alignment: input
[4,8,3,9]should predict target[8,3,9,2]. Keeping input and target identical would train copying instead of next-token prediction. - Data boundary: split the source text before making overlapping windows. Otherwise nearly identical windows can land on both sides of the train/validation boundary.
- Checkpoint evidence: include model config, model weights, optimizer state, step, tokenizer identity, data identity/split, seed policy, and validation evidence at minimum.
- Loss diagnosis: under a minimum-validation-loss rule, inspect the step-60 checkpoint first. Later training loss is not enough to override worsening held-out evidence.
- Generation: the model computes logits and therefore a next-token distribution. Temperature/top-k/top-p and the random draw belong to the sampling decision, not to retraining the model.
Pass condition
You are ready for 6.10 when you can trace the training target, checkpoint evidence, held-out selection rule, and generation step without future-token leakage or train/validation leakage. The next lessons change decoding behavior and evaluation—not the already trained weights unless explicitly stated.
Completion is stored locally on this device.