Level 10 project
Multimodal Tool-Using Assistant
Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.
Start here
Launch lesson: Tool-Using Assistant Integration Workshop
Prerequisite project: Evidence-Backed RAG
Goal
Build a bounded tool workflow whose authority and evidence remain inspectable:
request + trusted principal/state
→ tool selection
→ structured argument validation
→ authorization / approval
→ execution or direct response
→ normalized result
→ retry decision
→ state transition
→ final outcome
→ evaluation + release decisionTask
select_toolvalidate_argumentsnormalize_tool_resultauthorize_actionnext_retry_actiontransitionvalidate_run
Validation
Run these commands from the downloaded Project folder or the public materials repository root.
python projects/tests/l10/validate_submission.py \
projects/starters/l10/assistant.py \
path/to/your/tool-run.jsonRubric
Score each criterion from 0–4. A submission should score at least 16/20 and must not score 0 on authority/security or evaluation.
| Criterion | 4 — Strong evidence | 3 — Meets | 2 — Partial | 1 — Weak | 0 — Missing |
|---|---|---|---|---|---|
| Tool/schema design | Narrow task-shaped tools, typed schemas, extra-field policy, result allowlists, and clear read/write separation | Complete interfaces with minor gaps | Functional tools but loose schema or result boundaries | Broad ambiguous tools | No coherent tool interface |
| Authority, approval, and state | Trusted principal, least privilege, exact approval binding, explicit legal transitions, and denied-action traces | Correct boundaries with minor gaps | Some checks but one important authority/state boundary is weak | Prompt wording carries major policy responsibility | Unauthorized execution/approval bypass can pass |
| Error/retry/result handling | Failure categories, bounded retries, side-effect idempotency handling, normalized results, and provenance are demonstrated | Complete handling with minor gaps | Retry or normalization works but edge cases are shallow | Blind retries or raw result dumping | No reliable failure/result path |
| Multimodal and workflow evaluation | Text/image/audio slices, tool-selection/task metrics, zero-tolerance security metrics, and earliest-boundary failure analysis | Strong evaluation with small omissions | Fixed cases exist but slices/diagnosis are limited | Mostly anecdotal demos | No fixed evaluation |
| Reproducibility and communication | Versioned workflow/schema/policy/evaluator, commands, fixtures, limitations, optional live-path differences, and debug record are explicit | Reproducible with minor omissions | Several identities or limitations missing | Screenshots/output only | No reproducible package |
Objective checks
The validator checks tool selection, schema/range rejection, result allowlisting, role permissions, approval binding, retry/idempotency decisions, state transitions, input-evidence identity, text/image/audio slice metrics, run metrics, release recomputation, and intentional unsafe-fixture rejection.