본문으로 건너뛰기

Level 10 project

Multimodal Tool-Using Assistant

Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.

Launch lesson: Tool-Using Assistant Integration Workshop

Prerequisite project: Evidence-Backed RAG

Goal

Build a bounded tool workflow whose authority and evidence remain inspectable:

request + trusted principal/state
→ tool selection
→ structured argument validation
→ authorization / approval
→ execution or direct response
→ normalized result
→ retry decision
→ state transition
→ final outcome
→ evaluation + release decision

Task

  1. select_tool
  2. validate_arguments
  3. normalize_tool_result
  4. authorize_action
  5. next_retry_action
  6. transition
  7. validate_run

Validation

Run these commands from the downloaded Project folder or the public materials repository root.

python projects/tests/l10/validate_submission.py \
projects/starters/l10/assistant.py \
path/to/your/tool-run.json

Rubric

Score each criterion from 0–4. A submission should score at least 16/20 and must not score 0 on authority/security or evaluation.

Criterion4 — Strong evidence3 — Meets2 — Partial1 — Weak0 — Missing
Tool/schema designNarrow task-shaped tools, typed schemas, extra-field policy, result allowlists, and clear read/write separationComplete interfaces with minor gapsFunctional tools but loose schema or result boundariesBroad ambiguous toolsNo coherent tool interface
Authority, approval, and stateTrusted principal, least privilege, exact approval binding, explicit legal transitions, and denied-action tracesCorrect boundaries with minor gapsSome checks but one important authority/state boundary is weakPrompt wording carries major policy responsibilityUnauthorized execution/approval bypass can pass
Error/retry/result handlingFailure categories, bounded retries, side-effect idempotency handling, normalized results, and provenance are demonstratedComplete handling with minor gapsRetry or normalization works but edge cases are shallowBlind retries or raw result dumpingNo reliable failure/result path
Multimodal and workflow evaluationText/image/audio slices, tool-selection/task metrics, zero-tolerance security metrics, and earliest-boundary failure analysisStrong evaluation with small omissionsFixed cases exist but slices/diagnosis are limitedMostly anecdotal demosNo fixed evaluation
Reproducibility and communicationVersioned workflow/schema/policy/evaluator, commands, fixtures, limitations, optional live-path differences, and debug record are explicitReproducible with minor omissionsSeveral identities or limitations missingScreenshots/output onlyNo reproducible package

Objective checks

The validator checks tool selection, schema/range rejection, result allowlisting, role permissions, approval binding, retry/idempotency decisions, state transitions, input-evidence identity, text/image/audio slice metrics, run metrics, release recomputation, and intentional unsafe-fixture rejection.

← Back to all projects