Checkpoint — Evaluation and Safety Boundaries
Before moving into tool security, privacy, and red teaming, connect the first six Lessons into one evidence path.
Start with a release decision. You should be able to name the cases, expected behaviors, measures, slices, and critical blockers that support that decision. A single average is not enough when a smaller slice or zero-tolerance boundary can fail independently.
Explain how a golden set becomes a regression suite. Each important case needs an identity, provenance, observable expectation, and a reason it belongs. When an expectation changes, the change should be reviewed rather than silently edited alongside the candidate system.
Human evaluation should add evidence where deterministic checks are incomplete. Be able to define a rating dimension, provide observable anchors, preserve disagreement, and explain why reviewer judgment is informative without pretending it is perfect ground truth.
Then move from output quality to system safety. Draw the system boundary and identify assets, actors, tools, data stores, and trust boundaries. Use the risk context to decide which failures matter and where controls belong.
Finally, trace one prompt-injection path. Show how untrusted content reaches the model, what protected action or data it tries to influence, which trusted application control stops the path, and what evidence proves the control worked.
A strong answer connects all five ideas: threat paths become test cases; test cases enter regression suites; human review handles judgment-heavy properties; system controls enforce authority; release rules keep critical violations separate from averages.
Check yourself
- Why can a 99% aggregate pass rate still be insufficient for release?
- What information makes a golden case maintainable rather than just memorable?
- What does disagreement between careful human reviewers tell you to inspect?
- Why does the same model create different risk in a read-only assistant and a money-moving agent?
- Which part of a prompt-injection defense must remain outside the model?
Check your reasoning with an applied release case
Suppose a candidate passes 99% of a test suite, but the one failure is an approval bypass in a money-moving workflow.
The aggregate score does not erase that failure. If the release rule marks approval bypasses as zero-tolerance, the candidate fails that gate. Keep the failing case in the regression suite with stable identity, provenance, expected behavior, and the reason it is critical.
Then trace the safety boundary: untrusted content may influence model text, but trusted application code must still enforce authorization before money moves. Human reviewers can add judgment about unclear outputs, and disagreement is evidence to inspect—not a reason to average away a critical system failure.
The same model can have different risk in a read-only assistant and a money-moving agent because the assets, tools, permissions, and possible side effects differ.
Key Takeaways
- Evaluation begins with decisions and preserves the evidence needed to review them.
- Golden cases, human judgment, and system controls cover different kinds of uncertainty.
- Threat models connect assets and trust boundaries to concrete failure paths.
- Untrusted text cannot grant itself new authority.
- Critical failures should become regression evidence, not disappear into an aggregate score.
Next Lesson
Continue with L15.7 — Tool and Agent Security.
Completion is stored locally on this device.