Level 15 project
Ship, Evaluate, and Defend an AI Product
Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.
Start here
Launch lesson: Capstone Integration: Ship, Evaluate, and Defend
Prerequisite project: Production AI Service
Goal
Build the required release-evidence controller, then optionally run a separate learner-built real RAG pipeline.
Task
Complete the required TODO functions in capstone.py:
- regression and critical-case summaries;
- human-review mean and disagreement evidence;
- tool authorization using trusted principal state;
- retention/deletion eligibility;
- severity-aware threat-test summaries that preserve threat categories;
- dated frontier-claim validation;
- research-reproduction record validation;
- artifact license/provenance and governance-obligation evidence summaries, including missing required inventory entries;
- multi-dimensional release gates that include those evidence metrics;
- raw-evidence run validation, ISO-dated evidence checks, and release-decision recomputation.
The required acceptance path is deterministic and offline. It does not require a live model, provider API, network connection, GPU, cloud account, real user data, or secret.
Run:
python3 projects/tests/l15/validate_submission.py \
projects/starters/l15/capstone.py \
projects/tests/l15/fixtures/passing/capstone-run.jsonReference acceptance:
python3 projects/tests/l15/test_reference.pyLevel Labs:
python3 labs/notebooks/level-15/test_labs.pyValidation
Run these commands from the downloaded Project folder or the public materials repository root.
python3 projects/tests/l15/validate_submission.py \
projects/starters/l15/capstone.py \
projects/tests/l15/fixtures/passing/capstone-run.jsonpython3 projects/tests/l15/validate_product.py \
projects/starters/l15/product.pyRubric
Total: 100 points.
- Regression evaluation — 15: important cases, slices, and critical regressions are represented and recomputed from raw case outcomes.
- Human evaluation — 10: rating evidence preserves observable scores and meaningful disagreement rather than only a final average.
- Authorization and agent security — 15: trusted principal state controls side effects; generated identity claims cannot expand authority.
- Privacy, governance, and retention — 15: retention and hold logic is explicit, deterministic, and checked from raw record fields.
- Threat modeling and red-team evidence — 15: severity-aware adversarial tests preserve both blocked paths and critical control failures.
- Frontier and research evidence — 15: claims are scoped, dated, versioned, and paired with reproducible local evidence or explicit limitations.
- Release reasoning — 10: thresholds and zero-tolerance gates remain independent and the final release decision is recomputed.
- Reproducibility and communication — 5: offline acceptance, synthetic fixtures, environment metadata, and reviewable evidence are documented.
Full credit requires independent recomputation of critical metrics and release status. A high aggregate quality score cannot compensate for a zero-tolerance authorization, retention, or critical threat violation.