본문으로 건너뛰기
L15.13

Capstone Integration: Ship, Evaluate, and Defend

Goal

Integrate regression evaluation, human review, authorization, privacy/retention, threat testing, and scoped frontier claims into one release decision that can be independently recomputed.

Treat release as a checklist of independent results​

Imagine opening a school science exhibition. A project is not ready merely because the experiment worked once. The team checks the result and the safety rules, then checks who may operate the equipment and what data is displayed. It also checks what happens if something fails and verifies the checklist from the underlying records.

The capstone is that release checklist for an AI product. Each record answers a different question. Regression tests check known behavior. Human review covers judgment-heavy qualities. Authorization records check side effects, while privacy records check data handling. Red-team cases check hostile paths. Deployment identity checks whether the running version is the one that was reviewed. The final release decision should be recomputable from those pieces rather than trusted as one unexplained boolean.

Build the release record from raw results​

The final integration starts with a product candidate, not a model leaderboard. The candidate includes a model or models, prompts, retrieval, tools, policies, deployment configuration, data flows, and user experience. Shipping means accepting responsibility for how those parts behave together.

Build the release record from raw results. Keep regression case outcomes, human ratings, action attempts, policy decisions, retention records, red-team findings, deployment identity, and frontier/research records as separate inputs. The release summary may compress them, but it should not replace them.

regression cases ─┐
human review ─────┤
authorization ────┤
privacy/retention ├─→ recomputed release decision
threat tests ─────┤
deployment ID ────┘

Evaluate behavior and human judgment separately​

First evaluate product behavior. The regression suite should include common tasks, important slices, previous incidents, and critical boundaries. Compare the candidate with the current baseline by case ID so newly failing cases remain visible. An overall improvement does not excuse a critical regression unless the declared policy explicitly says it can.

Next interpret the reviewer results. Preserve the rubric and disagreement signal. A high average can coexist with a deeply split set of reviewers or with a separate critical factual failure. Human judgment contributes to the release record without becoming the sole source of truth.

Recompute authority and privacy decisions​

Then validate authority. Every side-effect attempt should identify the authenticated principal, proposed action, protected resource, relevant policy, approval state, and observed execution. The validator should detect an executed action that policy would have denied even if the final user-facing answer looked correct.

Privacy and governance checks should be equally mechanical. Verify that each record was accessed for an allowed purpose and that restricted data did not cross to an unauthorized destination. Then confirm that expired records followed the declared retention policy. The capstone uses synthetic records so you can test these rules without collecting real sensitive data.

Keep threat tests and frontier claims in scope​

Threat and red-team results contribute severity-aware failures. A blocked adversarial attempt shows that a particular control held for that case. An executed unauthorized path is a release blocker when the rule says zero critical security violations. Keep the attempted path and policy result so reviewers can understand which layer contained or failed to contain the threat.

Frontier claims are handled differently. A dated benchmark, protocol snapshot, or research reproduction may support a product hypothesis, but the release record must say what that source actually covers. A missing version, missing snapshot date, or conclusion stronger than the reproduced result becomes a record-quality problem rather than a hidden assumption.

Apply non-interchangeable release gates​

The release rule combines these dimensions without pretending they are interchangeable. A candidate may require minimum regression success and a minimum human-rating result. At the same time, it can require zero unauthorized executions, zero critical retention violations, and zero critical threat failures. Critical gates stay independent.

This is also where Level 14 operations return. A technically safe release still needs deployment identity, observability, rollback, and capacity controls. The capstone assumes those production foundations and adds the question: should this particular AI product version be allowed to receive wider traffic?

Validate with intentional failures​

The objective validator should recompute the decision from the raw fixtures. Do not trust release_passed, authorized, expired, or a precomputed risk count when primary fields are available. Independent recomputation is the common thread across the final Levels: raw records first, derived conclusion second.

Include one intentional failure fixture. Give it several distinct problems, for example:

  • a newly failing critical regression case;
  • an unauthorized executed action;
  • an expired restricted record kept without an allowed exception;
  • an unsupported frontier claim.

The validator should reject the run and identify the earliest violated release rule.

State residual risk and make the decision traceable​

A good capstone report also states uncertainty and residual risk. Passing the declared suite does not prove the system is universally safe. It means the candidate satisfied this reviewed set of requirements under the recorded tests, reviews, controls, and deployment environment. Future incidents, new capabilities, new users, or changed protocols can require new tests.

Finally, make the release decision explainable to someone who did not build the system. They should be able to trace one case from raw input to observed output. They should be able to follow one protected action from proposal to policy decision. They should also be able to follow one data record through retention logic and one frontier claim back to its dated source.

Predict

A candidate has higher regression and human-review scores, but one raw event shows an unauthorized side effect and the release rule allows zero. What should the capstone decide?

Run the Docker-environment Lab preflight​

Start with the passing capstone fixture:

python3 labs/notebooks/level-15/l15-13-capstone-check.py projects/reference/l15/sample-run/capstone-run.json

Confirm it reaches PASS: Level 15 integration trace.

Before running the failure fixture, inspect only action act-2: the trusted principal role is support, the proposal asks for refund_order, and executed is true. Predict the recomputed unauthorized_actions count and whether a release rule that allows zero should pass.

Then run:

python3 labs/notebooks/level-15/l15-13-capstone-check.py projects/reference/l15/failure-run/capstone-run.json

Confirm the checker rejects the integration trace. Then continue tracing the remaining raw regression, retention, red-team, frontier-claim, and reproduction records instead of trusting the stored metrics or release_passed flag.

Loading lab…

Quick Check

1. Why keep the raw records after computing a release summary?
2. How should a blocked red-team attempt be interpreted?
3. What does a passing capstone release record establish?

0 of 3 questions answered.

Explain it back​

Walk through one candidate release. Name the regression results, human-review results, authorization record, retention record, threat-test result, and frontier/research source. State which failures are thresholds and which are zero-tolerance gates.

Key Takeaways

  • Ship the whole product candidate, so evaluate the whole product boundary.
  • Preserve raw regression, human, authorization, governance, threat, and frontier evidence.
  • Recompute critical derived metrics and release decisions independently.
  • Keep zero-tolerance safety/security gates separate from aggregate quality scores.
  • A passing release is scoped evidence, not a claim of universal or permanent safety.

Next Lesson

You are ready for the Level 15 project, Ship, Evaluate, and Defend an AI Product. Completing the project produces the evidence package for final course QA and acceptance; it does not by itself declare the program accepted.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Ship, Evaluate, and Defend an AI Product

View progress