Multimodal Evaluation
Goal
Evaluate a multimodal assistant by checking the image or audio input stage, the model/tool stage, and the final answer separately, then combine those checks without hiding modality-specific failures.
Start with one wrong receipt answer
Imagine an assistant receives a photo of a receipt and answers:
The total is $42.
The final answer is wrong. But there are several different reasons it could be wrong:
receipt image
↓
text/number extraction
↓
reasoning or tool use
↓
final answer
Maybe the image stage read $24 as $42. Maybe the image was read correctly but the calculator tool received the wrong field. Maybe every intermediate value was correct and the final response copied the wrong number.
One overall accuracy score cannot tell those failures apart. A useful multimodal evaluation marks what each case actually requires—text, image content, audio content, a tool result, or some combination—and records the first stage that went wrong.
For example, a receipt case may expect the image stage to recover the merchant and total, then expect the answer stage to use those recovered values correctly. A speech case may separately check transcription, intent, and the final action. The point is not to create more scores for decoration. It is to make the failure location visible enough that you know what to fix.
Separate component and end-to-end metrics
Useful slices include:
- tool-selection accuracy;
- argument-validation pass/fail correctness;
- visual observation accuracy;
- transcription word/field accuracy;
- grounded final-answer accuracy;
- unauthorized-action count;
- unnecessary-tool-call rate;
- completion/abstention rate.
Component metrics explain why an end-to-end case failed.
Evaluate paired perturbations
Strong tests change one thing while keeping the task fixed:
clear image → blurred image
clean speech → noisy speech
valid tool result → timeout
authorized order → different-tenant order
The expected behavior should change only where the evidence or permission change requires it.
Do not let text dominate the benchmark
If 95% of cases are easy text questions and 5% require image reasoning, a high overall score may say little about visual capability.
Report modality counts and per-slice results. Weighting should match the product risk and use distribution, not convenience.
Record failure stage
For each failed case, identify the earliest stage:
input/preprocessing
perception/transcription
tool selection
argument validation
execution
result normalization
final reasoning
security/approval
That label makes the next experiment targeted.
Define success at the case level before averaging
Each evaluation case should say what counts as success: which evidence must be used, which tool is allowed, whether abstention is acceptable, and what security constraints are non-negotiable. This prevents a grader from moving the goalposts after seeing a fluent output.
For example, a blurred receipt image may be intentionally unanswerable. If the expected behavior is to request a clearer image, an abstention is a success rather than a failure. The same case becomes dangerous if the assistant guesses an amount and then calls a payment tool.
Security metrics are not ordinary quality trade-offs
Unauthorized execution, approval bypass, secret exposure, or sandbox escape should usually be reported as separate counts with strict thresholds. Averaging them into a generic score can make one serious failure disappear behind many easy correct cases. Release rules should reflect the asymmetry between harmless wording mistakes and boundary violations.
Predict
Run the local Lab
python labs/notebooks/level-10/l10-11-multimodal-eval.py
- Run the command. The report shows
overall: 0.75,by_modality: text 1.0, image 0.0, audio 1.0, andsecurity_failures: 0. - Add two easy text cases. After the
a1line incases, add:
{"id": "t3", "modality": "text", "passed": True, "security_failure": False},
{"id": "t4", "modality": "text", "passed": True, "security_failure": False},
- Rerun.
overallrises to about0.83, butimageis still0.0. The average improved only because easy cases were added. The script then stops with anAssertionError, which is expected: its final checks were written for the original values. Change the value back afterward. - Remove the two lines. Now change
t2to"security_failure": Trueand rerun.overallstays0.75, butsecurity_failuresbecomes1. A zero-tolerance count must be reported on its own, because the average does not move at all.
Loading lab…
Quick Check
Key Takeaways
- Evaluation should mirror the multimodal/tool pipeline.
- Label required evidence and allowed tools per case.
- Keep component metrics and end-to-end metrics together.
- Use paired perturbations to isolate failure causes.
- Report modality slices and security failures separately.
Next Lesson
Next, put permissions, approvals, untrusted inputs, and tool authority into one explicit security model.
References
- Liang et al., Holistic Evaluation of Language Models.
- Radford et al., CLIP.
- Radford et al., Whisper.
Completion is stored locally on this device.