Tiny LLM Integration Workshop
Goal
Verify the complete tiny-language-model evidence chain across data, model, training/checkpoint, evaluation, failure diagnosis, and reproducible packaging.
Individual pieces can pass while their interfaces disagree. The final workshop treats every handoff as a contract:
text → tokenizer → windows/targets → model → loss/update → checkpoint
↓
held-out loss + fixed prompts/samples
↓
versioned run package
The workshop is successful when another reviewer can follow that chain and understand both a passing run and an intentional failure.
Follow one failure upstream instead of patching the end
Suppose generation suddenly becomes nonsense after a change. The final text is only the last visible symptom. Trace backward:
sample
← decoding settings
← logits/checkpoint
← model configuration
← token IDs
← tokenizer/vocabulary
If checkpoint and vocabulary fingerprints disagree, changing temperature is downstream of the real problem. The token IDs may already mean the wrong things before the model sees them.
The same reasoning applies to evaluation. A perfectly valid metrics.json file can still be attached to the wrong checkpoint. Integration correctness therefore requires identity relationships, not only individually well-formed files.
For the workshop, make each invariant answer a concrete question: “How do I know this tokenizer belongs with this checkpoint?” “How do I know validation text was not used for updates?” “How do I know the reported sample used the recorded decoding configuration?” The reduced validator is useful because it checks these relationships without requiring a full retraining run.
Integration tests check relationships across components
Unit tests can prove that the tokenizer encodes text or that the model returns an expected shape.
An integration test asks whether the pieces agree with one another:
tokenizer vocabulary size == model vocab_size
checkpoint config == constructed model config
evaluation checkpoint ID == reported metrics checkpoint ID
prompt token count <= model context limit
These are cross-component contracts. A failure here may not belong to any one component in isolation.
Trace one sample all the way through
Take one fixed prompt and record:
- raw text;
- token IDs;
- input shape;
- checkpoint identity;
- logits shape at the final position;
- decoding configuration;
- selected/generated IDs;
- decoded output.
If the final output is surprising, this trace gives you boundaries to inspect instead of restarting the whole training run.
When many individually correct parts form one system, verify the relationships between them explicitly.
Predict
Write invariants before running
Record the expected conditions:
- token IDs lie in the configured vocabulary;
- input/target shifts are correct and train/validation regions are separate;
- model dimensions and checkpoint config match;
- loss is finite and checkpoint selection is traceable;
- tokenizer/checkpoint identities agree;
- evaluation uses fixed held-out data/prompts/decoding configuration;
- run metadata is complete enough for the reduced validator.
Run the passing reference package first:
python labs/notebooks/level-06/l06-14-integration-check.py projects/reference/l06/sample-run
It should print PASS: integration invariants hold for p06-reference-reduced.
Before running the intentional mismatch, inspect only the two identity fields in the failure fixture and predict which integration invariant should fail while the package itself can still be structurally complete. Then run:
python labs/notebooks/level-06/l06-14-integration-check.py projects/reference/l06/failure-run
Confirm the failure is the tokenizer/checkpoint fingerprint relationship rather than a missing-file error. The checker accepts either a run directory or the corresponding run.json path.
Loading lab…
Fix the producing artifact pairing rather than weakening the validator.
The deterministic reduced path is important: reviewers can verify integration and packaging contracts without repeating a long training session or having the same hardware.
Quick Check
Final explain-back
Explain the complete workflow to a learner who finished Level 5 but has never trained a language model. Name at least four distinct invariants—data, causal/shape, checkpoint, and evaluation/package—and give one example of how a valid-looking downstream artifact could still be attached to the wrong upstream state.
Key Takeaways
- End-to-end correctness depends on contracts between data, model, training, evaluation, and packaging.
- Tokenizer/config/checkpoint identities must stay aligned.
- Fixed evaluation conditions make comparisons interpretable.
- Reduced deterministic validation makes technical review practical.
- Fix inconsistent source evidence instead of weakening its checks.
Next Lesson
You are ready for Train and Ship a Tiny LLM. After the project, the curriculum moves from building the model to using modern LLMs reliably as probabilistic components.
References
- Radford et al., Language Models are Unsupervised Multitask Learners.
Completion is stored locally on this device.
Level project unlocked: Train and Ship a Tiny LLM