본문으로 건너뛰기
L10.11

Multimodal Evaluation

Goal

Evaluate a multimodal assistant by checking the image or audio input stage, the model/tool stage, and the final answer separately, then combine those checks without hiding modality-specific failures.

Start with one wrong receipt answer​

Imagine an assistant receives a photo of a receipt and answers:

The total is $42.

The final answer is wrong. But there are several different reasons it could be wrong:

receipt image
↓
text/number extraction
↓
reasoning or tool use
↓
final answer

Maybe the image stage read $24 as $42. Maybe the image was read correctly but the calculator tool received the wrong field. Maybe every intermediate value was correct and the final response copied the wrong number.

One overall accuracy score cannot tell those failures apart. A useful multimodal evaluation marks what each case actually requires—text, image content, audio content, a tool result, or some combination—and records the first stage that went wrong.

For example, a receipt case may expect the image stage to recover the merchant and total, then expect the answer stage to use those recovered values correctly. A speech case may separately check transcription, intent, and the final action. The point is not to create more scores for decoration. It is to make the failure location visible enough that you know what to fix.

Separate component and end-to-end metrics​

Useful slices include:

  • tool-selection accuracy;
  • argument-validation pass/fail correctness;
  • visual observation accuracy;
  • transcription word/field accuracy;
  • grounded final-answer accuracy;
  • unauthorized-action count;
  • unnecessary-tool-call rate;
  • completion/abstention rate.

Component metrics explain why an end-to-end case failed.

Evaluate paired perturbations​

Strong tests change one thing while keeping the task fixed:

clear image → blurred image
clean speech → noisy speech
valid tool result → timeout
authorized order → different-tenant order

The expected behavior should change only where the evidence or permission change requires it.

Do not let text dominate the benchmark​

If 95% of cases are easy text questions and 5% require image reasoning, a high overall score may say little about visual capability.

Report modality counts and per-slice results. Weighting should match the product risk and use distribution, not convenience.

Record failure stage​

For each failed case, identify the earliest stage:

input/preprocessing
perception/transcription
tool selection
argument validation
execution
result normalization
final reasoning
security/approval

That label makes the next experiment targeted.

Define success at the case level before averaging​

Each evaluation case should say what counts as success: which evidence must be used, which tool is allowed, whether abstention is acceptable, and what security constraints are non-negotiable. This prevents a grader from moving the goalposts after seeing a fluent output.

For example, a blurred receipt image may be intentionally unanswerable. If the expected behavior is to request a clearer image, an abstention is a success rather than a failure. The same case becomes dangerous if the assistant guesses an amount and then calls a payment tool.

Security metrics are not ordinary quality trade-offs​

Unauthorized execution, approval bypass, secret exposure, or sandbox escape should usually be reported as separate counts with strict thresholds. Averaging them into a generic score can make one serious failure disappear behind many easy correct cases. Release rules should reflect the asymmetry between harmless wording mistakes and boundary violations.

Predict

Overall accuracy rises because many easy text cases were added, while image accuracy is unchanged. What conclusion is justified?

Run the local Lab​

python labs/notebooks/level-10/l10-11-multimodal-eval.py
  1. Run the command. The report shows overall: 0.75, by_modality: text 1.0, image 0.0, audio 1.0, and security_failures: 0.
  2. Add two easy text cases. After the a1 line in cases, add:
{"id": "t3", "modality": "text", "passed": True, "security_failure": False},
{"id": "t4", "modality": "text", "passed": True, "security_failure": False},
  1. Rerun. overall rises to about 0.83, but image is still 0.0. The average improved only because easy cases were added. The script then stops with an AssertionError, which is expected: its final checks were written for the original values. Change the value back afterward.
  2. Remove the two lines. Now change t2 to "security_failure": True and rerun. overall stays 0.75, but security_failures becomes 1. A zero-tolerance count must be reported on its own, because the average does not move at all.

Loading lab…

Quick Check

1. Why keep component metrics beside end-to-end accuracy?
2. What is a useful paired perturbation?
3. Why report modality counts?

0 of 3 questions answered.

Key Takeaways

  • Evaluation should mirror the multimodal/tool pipeline.
  • Label required evidence and allowed tools per case.
  • Keep component metrics and end-to-end metrics together.
  • Use paired perturbations to isolate failure causes.
  • Report modality slices and security failures separately.

Next Lesson

Next, put permissions, approvals, untrusted inputs, and tool authority into one explicit security model.

References

Lesson actions

Completion is stored locally on this device.

View progress