Skip to main content
L13.11

Multi-Agent Evaluation

Goal

Measure whether a multi-agent system succeeds, sends work to the right agent, preserves evidence during handoffs, avoids unsafe authority, and uses coordination efficiently.

Evaluate the team, not only each agent​

Think about a relay race. Every runner can be individually fast, but the team can still lose if the baton is handed to the wrong person, dropped during an exchange, or carried around the track twice. Multi-agent systems have the same kind of team-level failure: local performance can look good while coordination is wrong.

For example, suppose a coordinator creates four subtasks. Three go to agents with the right advertised skill, but one finance task goes to a travel agent. Even if all four agents return something, routing accuracy is only 3 / 4 = 75%. Now suppose one correctly chosen child never receives the account ID it needs. That is a separate handoff completeness failure. Measuring only "did each process finish?" would hide both problems.

So evaluate the whole path, not only the final answer. Record the parent task, child tasks, selected agents, messages, artifacts, shared-state updates, retries, conflicts, and final outcome. These links show where the failure happened.

Measure routing and handoff quality​

Keep task success, but add two simple measures. Routing accuracy asks whether the coordinator chose an agent whose advertised skill matched the subtask. Handoff completeness asks whether the child received the needed inputs and returned the required evidence. For example, a finance specialist that never receives the account ID cannot complete a correct handoff.

routing accuracy: 3 / 4 = 75%
handoff completeness: 3 / 4 = 75%
unauthorized calls: 0 required

Shared state can create conflicts. Count stale-write rejections, unresolved conflicts, and cases where weaker evidence replaced stronger evidence. Many detected conflicts are not automatically bad. What matters is whether the system resolves them safely.

Track duplicate work and authority failures​

Also track duplicate work. Two agents may run the same expensive lookup because the coordinator failed to notice that the subtasks were equivalent. Stable task and operation IDs help detect that waste.

Some authority failures should be zero tolerance. One unauthorized remote call should not disappear inside a high average success rate. The same rule can apply to a credential-scope violation or an artifact accepted from the wrong task.

Coordination also has a cost. Count messages, child tasks, latency, and compute. If one design needs 30 handoffs and another needs 5 for the same reliable result, that difference matters.

Replay failures with fixtures and simulations​

Use fixed fixtures for exact safety rules. Then use seeded simulations for changing conditions such as agent availability or conflicting observations. Keep enough replay data to inspect a failed run later.

AgentBench provides a useful precedent for evaluating interactive agent behavior across environments. In this level, extend that idea to the coordination layer: evaluate the path between agents, not just the final answer.

Predict

Every child agent completes its local task, but the parent routes a finance request to a travel specialist and then discards the returned evidence. How should system evaluation score this?

Run the Docker-environment Lab preflight​

This activity is registered for the Docker-oriented environment used by multi-agent operational work. Start with the deterministic Python preflight so protocol/controller failures remain distinguishable from container or network setup problems.

Run:

python3 labs/notebooks/level-13/l13-11-evaluation.py

The Lab recomputes routing and coordination metrics from raw traces rather than trusting stored summaries.

  1. Run it unchanged. Record the routing accuracy and note that t3 is already a deliberate mismatch.
  2. In trace t2, find "assigned": "policy-agent". Before editing, predict which metric should fall if only that assignment changes to an agent without policy_lookup.
  3. Change only t2's assigned agent to "travel-agent", then rerun.
  4. Confirm routing accuracy falls because two traces are now capability-mismatched, while the unchanged handoff and authorization fields keep their own metrics separate.

Loading lab…

Quick Check

1. Why is final task success insufficient for multi-agent evaluation?
2. What should handoff completeness measure?
3. How should unauthorized remote calls enter release rules?

0 of 3 questions answered.

Explain it back​

Design a multi-agent scorecard with at least six metrics. Separate average performance metrics from zero-tolerance trust or authority controls.

Key Takeaways

  • Evaluate the coordination trajectory, not just individual agents.
  • Measure routing and handoff quality.
  • Track conflict resolution and duplicate work.
  • Coordination cost includes messages, child tasks, latency, and compute.
  • Critical authority violations remain independent release blockers.

Next Lesson

Next, L13.12 — Interoperable Agent System Workshop integrates MCP, A2A, trust, shared state, and evaluation.

References

Lesson actions

Completion is stored locally on this device.

View progress