Skip to main content
L13.12

Interoperable Agent System Workshop

Goal

Combine MCP and A2A boundaries with trust rules, task IDs, conflict handling, and release evidence.

Follow one user goal across two protocol boundaries​

Start with one ordinary request: "Check the order status, then ask the shipping specialist whether the delay needs action." The harness may use MCP for the order lookup. It may use an A2A agent for the specialist task. Those are two different boundaries inside one user goal.

The order lookup needs an operation identity so a retry can be recognized. The delegated specialist work needs its own child task ID. Later status updates and artifacts can then be attached to the correct job. If the specialist returns a file, the parent must know which agent and child task produced it before trusting the file. These IDs answer different recovery and trust questions.

Check MCP and A2A separately​

The MCP side is pinned to 2026-07-28. Before a tool can run, the harness checks version, endpoint trust, discovered capability, argument shape, and local authorization. A protocol-valid request can still fail local policy.

The A2A side is pinned to v1.0.0. The harness checks the selected Agent Card, tracks the remote task ID, and accepts messages or artifacts only for the expected task. The child also gets less authority than the parent.

Keep different IDs for different jobs​

Keep different IDs for different jobs. A parent task can contain an MCP operation ID and an A2A child task ID. One generic ID cannot replace all of them because they support different recovery and audit questions.

parent task: P-7
MCP operation: op-22
A2A child task: child-9
artifact source: agent=shipping-specialist, task=child-9

Protect shared state and provenance​

Shared-state updates use version checks. A child artifact may add evidence, but only the assigned owner may move the parent into a final state. If two updates race, record a conflict instead of silently accepting the last write.

Remote outputs also keep provenance, meaning where they came from. An MCP result records its endpoint and capability. An A2A artifact records its agent and task. Later policy can use that source information.

Use the same trust rule on both protocols: narrow credentials and data, verify endpoint or agent identity, and never let successful parsing bypass authorization.

Recompute release evidence from raw events​

Recompute evaluation metrics from raw events. Track protocol and trust failures such as version mismatches, unauthorized calls, wrong-task artifacts, and state conflicts. Track system quality too: routing accuracy, duplicate operations, task success, message count, latency, and cost.

Separate normal quality thresholds from critical controls. A small increase in message count might be acceptable. An unauthorized action or an artifact from the wrong task may block release.

When debugging, find the earliest broken boundary. If the system accepted an artifact from another task, changing the final prompt does not repair that source-identity bug.

Predict

A valid A2A artifact arrives with a task ID that does not match the delegated child task. What should the parent do?

Run the Docker-environment Lab preflight​

This activity is registered for the Docker-oriented environment used by multi-agent operational work. Start with the deterministic Python preflight so protocol/controller failures remain distinguishable from container or network setup problems.

Start with the passing interoperable-system fixture:

python3 labs/notebooks/level-13/l13-12-integration-check.py projects/reference/l13/sample-run/system-run.json

Confirm the validator recomputes a passing release from the raw events.

For a one-variable experiment, copy the passing fixture to a temporary file, then change only the A2A artifact's task ID from "task-7" to "task-other". Before rerunning, predict which metric and release gate should change while protocol versions and MCP authorization remain fixed.

cp projects/reference/l13/sample-run/system-run.json /tmp/l13-system-run.json
python3 labs/notebooks/level-13/l13-12-integration-check.py /tmp/l13-system-run.json

After editing the copied JSON, rerun the second command. The validator should report one wrong-task artifact and fail the corresponding release gate.

Then run the repository's multi-failure fixture as a debugging challenge:

python3 labs/notebooks/level-13/l13-12-integration-check.py projects/reference/l13/failure-run/system-run.json

That fixture intentionally breaks several boundaries. Locate the first broken invariant in each affected run instead of treating the stored release_passed flag as evidence.

Loading lab…

Quick Check

1. What should happen before an MCP capability becomes executable?
2. What is the purpose of preserving separate parent, operation, and child-task IDs?
3. How should system release metrics be computed?

0 of 3 questions answered.

Explain it back​

Walk through one task that reads an MCP resource, calls one MCP tool, delegates one A2A child task, receives an artifact, resolves one shared-state update, and produces release evidence.

Key Takeaways

  • Pin and validate the protocol versions used by each boundary.
  • Keep remote capability and artifact provenance explicit.
  • Local trust and authorization remain authoritative.
  • Use versioned shared state and distinct operation/task identities.
  • Recompute release metrics from raw events across the full system.

Next Lesson

You are ready for the Level 13 project: build and validate the Interoperable Agent System.

Bridge to Level 14: Level 13 makes separate tools and agents communicate through explicit, testable boundaries. The next problem is keeping that whole system reliable when real requests arrive at the same time. Level 14 therefore moves from interoperability to production serving: APIs, batching, memory limits, deployment identity, observability, rollouts, and recovery.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Interoperable Agent System

View progress