Skip to main content

Level 6 project

Train and Ship a Tiny LLM

Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.

Launch lesson: Tiny LLM Integration Workshop

Prerequisite project: Build a Mini Transformer

Goal

Train or reuse the validated Level 6 tiny decoder-only language-model path, evaluate it on fixed evidence, diagnose one failure, and ship a reproducible run package that another learner can inspect.

Task

Complete the TODOs in run_package.py and produce a run.json from your training/evaluation evidence.

Your implementation must provide:

  1. perplexity(loss) for positive finite average natural-log loss;
  2. repeated_bigram_rate(text) as a simple generation-failure indicator;
  3. choose_best(records) selecting the lowest validation-loss record;
  4. validate_run(data) rejecting missing or inconsistent reproducibility evidence.

Your run.json must include:

  • run_id and integer seed;
  • model config including vocabulary/context/model dimensions;
  • tokenizer_fingerprint;
  • checkpoint_fingerprint and matching checkpoint_tokenizer_fingerprint;
  • train/validation loss history or selected validation loss evidence;
  • at least two fixed evaluation samples with their prompt/decoding settings;
  • environment with Python/runtime information;
  • debug_record containing failure, evidence, hypothesis, focused fix, and result;
  • a limitations statement.

Validation

Run these commands from the downloaded Project folder or the public materials repository root.

python projects/tests/l06/validate_submission.py \
projects/starters/l06/run_package.py \
path/to/your/run.json
uv run python projects/tests/l06/validate_submission.py \
projects/starters/l06/run_package.py \
path/to/your/run.json

Rubric

Score each criterion from 0–4. A submission should score at least 16/20 overall and must not score 0 on reproducibility or debugging.

Criterion4 — Strong evidence3 — Meets2 — Partial1 — Weak0 — Missing
End-to-end language-model workflowCorrect next-token/data/model/training story; training evidence is finite and tied to a decoder-only configuration; selected checkpoint is justifiedComplete workflow with minor gapsSome stages evidenced but important connection missingMostly description without run evidenceNo credible training workflow
Evaluation and controllable samplingHeld-out loss plus fixed prompt samples; decoding settings recorded; controlled comparison changes one factor; limitations statedUses held-out metric and fixed samples with settingsEvaluation present but comparison/provenance weakOne cherry-picked sample or train loss onlyNo evaluation
Failure analysis and debuggingReproduces a concrete failure, separates evidence/hypothesis, makes one focused fix/check, records result and next implicationComplete failure/debug traceFailure shown but cause/evidence reasoning incompleteVague “fixed it” noteNo failure/debug path
Reproducibility and packagingSeed, config, tokenizer/checkpoint fingerprints, environment, commands, artifact provenance, metrics, samples, and reduced validation path all align; no secretsPackage can be restored/reviewed with small omissionsImportant metadata missing or ambiguousMostly an unlabeled artifactMissing package or secrets committed
Communication and engineering judgmentExplains tensor/data contracts, resource expectations, what evidence proves, what it does not prove, and next testClear report with sensible tradeoffsUnderstandable but shallow tradeoff reasoningMostly output dumpNo explanation

Exit-skill mapping

  • Train a tiny decoder-only LM end to end: End-to-end workflow.
  • Diagnose loss/generation failures: Failure analysis and debugging; evaluation.
  • Sample controllably: Evaluation and controllable sampling.
  • Package a reproducible run: Reproducibility and packaging.

Objective checks

The repository validator verifies the reduced-path functions, required run metadata, compatible tokenizer/checkpoint fingerprints, fixed sample records, valid configuration shapes, and intentional failure rejection. Human review scores whether the actual training/evaluation evidence and explanations are technically convincing.

Assessment guardrail

Generated text does not need to match the reference sample. Tiny model quality is intentionally limited. The rubric rewards controlled evidence, correct contracts, debugging, and reproducibility rather than fluent prose alone.

Want the from-scratch version?

This project trains and packages a working tiny LLM. Want to build every part from scratch against tests and debug planted failures? Tiny LLM Builder

← Back to all projects