Agent Test Fixtures
Goal
Build deterministic test fixtures for agent trajectories so controller behavior can be checked repeatedly without relying on a live model or unstable external service.
Turn agent behavior into known cases
Agent systems can appear difficult to test because model outputs vary and tools may depend on external services. The solution is not to give up on deterministic tests. Instead, test the harness with fixtures: fixed inputs, fixed model proposals, fixed tool outcomes, and expected controller behavior.
A fixture is a known scenario. For example, one fixture may describe a successful order lookup. Another may describe an expired approval, a duplicate side-effect delivery, or a child task that exceeds its budget. The harness runs against these fixed events and should produce the expected transitions and metrics.
Fixtures are especially useful for boundaries the model does not own. Permission checks, state transitions, retry classification, idempotency handling, memory policy, queue recovery, and release rules should be deterministic. A live model is not required to prove that an unauthorized refund is rejected.
{"case_id":"expired-approval","role":"support","action":"refund","approval_expired":true,"expected":"deny"}
Make fixtures recomputable
Good fixtures contain enough information to reproduce the decision. A tool failure fixture should say which attempt failed, what error class occurred, and what operation ID was used. A retry test that only says 'something went wrong' cannot verify precise recovery behavior.
Negative fixtures are as important as passing fixtures. A test suite should intentionally include stale approvals, unauthorized tools, malformed child results, duplicate messages, and continued execution after a stop condition. The test should fail if the harness incorrectly accepts those cases.
Do not encode the answer only in a trusted boolean inside the fixture. If a fixture says authorized:false, a strong validator should recompute authorization from role and policy rather than trusting that label. This prevents the test data from accidentally grading itself.
Use small failures to protect refactors
Fixtures also protect refactoring. You may replace a queue implementation or reorganize controller classes while expecting the same externally visible behavior. If the fixtures still pass, the refactor preserved the tested contract.
Keep fixture scope understandable. One giant scenario with thirty intertwined failures is hard to debug. Prefer small cases that isolate one boundary, plus a few end-to-end integration fixtures that test interactions between boundaries.
Later, simulation will vary inputs across many runs. Fixtures come first because they establish exact known cases. They answer 'does this specific safety and recovery rule work?' before simulation asks 'how often does the system remain reliable across a range of conditions?'
Predict
Run the Docker-environment Lab preflight
The full activity is registered for the Docker-oriented environment used by later operational work. Start with this deterministic Python preflight from your repository checkout so you can separate controller-logic failures from Docker, cloud, GPU, or credential setup problems.
Run:
python labs/notebooks/level-12/l12-09-test-fixtures.py
The Lab runs one passing and several failing deterministic fixtures. First inspect which rule rejects each failure. Then flip the derived authorized field in the unauthorized fixture and rerun. The validator should still reject it because policy is recomputed from raw role and action data.
Loading lab…
Quick Check
Explain it back
Describe three fixtures for the same side-effecting tool: one success, one retry-safe timeout, and one unauthorized attempt. State what the validator should recompute in each case.
Key Takeaways
- Deterministic fixtures make harness behavior repeatable.
- Test controller boundaries without requiring a live model.
- Include negative cases for important failures.
- Recompute critical derived facts from raw evidence.
- Use small focused fixtures plus a few integration scenarios.
Next Lesson
Next, L12.10 — Simulation-Based Evaluation expands exact fixtures into controlled families of delays, failures, and user behavior.
References
- Liu et al., AgentBench.
Completion is stored locally on this device.