Skip to main content
L12.13

Agent Harness Integration Workshop

Goal

Integrate durable tasks, replay-safe effects, tracing, permissions, memory policy, bounded parallelism, delegation, and evaluation into one testable agent harness.

Follow one task end to end​

Follow one task through the harness before thinking about every component at once. A request enters as a durable task record. A worker claims it. The controller chooses and validates an action. The tool adapter performs bounded work, and the result is recorded before the task continues. If the worker crashes, the next worker should recover from durable records instead of guessing from an incomplete conversation.

Now add the harder cases one boundary at a time. A retry keeps the same operation identity. A parallel read keeps its own child trace. A delegated subtask gets narrower authority. A memory summary cannot overwrite structured approval state. The workshop is one connected system, but each safety property should still be explainable at the boundary that enforces it.

The task begins in durable state. A queue delivers the task to a worker, which claims it under an ownership rule. The controller loads the current task record rather than assuming a fresh start. This lets a restarted worker continue from known state.

durable task → claim → controller → policy → tool adapter → durable result
↘ trace event at each boundary

Make effects replay-safe and observable​

Every side-effecting operation receives a stable operation ID. Before retrying an uncertain effect, the harness checks completion evidence or reconciles with the external adapter. This prevents a worker restart or duplicate task delivery from becoming a duplicate refund, message, or write.

Each important boundary emits structured trace events. The workshop records task state changes, selected actions, policy results, tool attempts, retry reasons, child-task creation, and terminal outcomes. These records are observable execution evidence, not hidden reasoning.

Keep authority boundaries separate​

Permission and isolation checks remain independent. The policy decides whether an action is allowed for the current principal and task. The execution profile restricts accessible files, endpoints, credentials, and resources. A permitted action can still be rejected if its requested execution environment exceeds the allowed profile.

Memory retrieval uses scope, provenance, freshness, and bounded result count. Authority fields such as approvals and operation completion stay in structured state rather than being trusted from a summary.

Coordinate parallel and delegated work​

The harness may run independent read tools in parallel. It must still enforce a concurrency limit, preserve child spans, and define how partial failures are joined. Conflicting writes remain serialized unless an explicit coordination mechanism makes them safe.

Delegated subtasks receive narrower tools and budgets than the parent. A child result must return required evidence fields. The parent validates that result before using it in a later decision.

Evaluate and debug the earliest failure​

Evaluation runs deterministic fixtures first, then seeded simulations. The release decision combines several ordinary metrics: task success, action-selection accuracy, recovery behavior, latency, and cost. It keeps zero-tolerance controls separate for unauthorized execution, approval bypass, duplicate side effects, and uncontrolled runaway work.

The final habit is to debug the earliest broken boundary. If a release fails because a duplicate effect occurred, do not start by rewriting prompts. Inspect whether operation identity, durable completion, reconciliation, or retry policy failed first. Harness engineering is about making those boundaries explicit enough to test.

Predict

A release fixture shows a duplicate refund after a worker crash. Which boundary should you inspect first?

Run the Docker-environment Lab preflight​

The full activity is registered for the Docker-oriented environment used by later operational work. Start with this deterministic Python preflight from your repository checkout so you can separate controller-logic failures from Docker, cloud, GPU, or credential setup problems.

Run the passing fixture:

python labs/notebooks/level-12/l12-13-integration-check.py projects/reference/l12/sample-run/harness-run.json

Then run the intentional failure fixture:

python labs/notebooks/level-12/l12-13-integration-check.py projects/reference/l12/failure-run/harness-run.json

The checker recomputes the release evidence from raw traces instead of trusting the stored summary. The passing fixture should print a PASS: marker; the failure fixture should exit non-zero and identify a violated release rule.

Loading lab…

Quick Check

1. What is the strongest purpose of the integrated harness?
2. What should happen before retrying an uncertain side effect?
3. How should a failed release be debugged?

0 of 3 questions answered.

Explain it back​

Walk through one task from queue delivery to terminal state. Include one retry, one parallel read, one delegated subtask, one permission check, one trace relationship, and the release evidence produced at the end.

Key Takeaways

  • Durable state and stable operation identity make recovery safer.
  • Tracing connects task, retry, delegation, and terminal evidence.
  • Permissions, isolation, and memory policy remain separate control layers.
  • Parallelism and delegation must stay bounded.
  • Release decisions should recompute critical metrics from raw evidence.

Next Lesson

You are ready for the Level 12 project: build and validate the Production-Grade Agent Harness before moving to interoperability in Level 13.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Production-Grade Agent Harness

View progress