본문으로 건너뛰기

Level 11 project

Bounded Agent

Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.

Launch lesson: Bounded Agent Integration Workshop

Prerequisite project: Multimodal Tool-Using Assistant

Goal

Build a small agent controller whose loop is explicit, bounded, reviewable, and testable.

Task

Open projects/starters/l11/agent.py and complete:

  1. choose_subgoal
  2. select_action
  3. update_working_memory
  4. memory_write_allowed
  5. approval_valid
  6. stop_reason
  7. transition
  8. validate_run

Validation

Run these commands from the downloaded Project folder or the public materials repository root.

python projects/tests/l11/validate_submission.py   projects/starters/l11/agent.py   projects/tests/l11/fixtures/passing/agent-run.json

Rubric

Score each criterion from 0–4. A submission should score at least 16/20 and must not score 0 on authority/boundedness or evaluation.

Criterion4 — Strong evidence3 — Meets2 — Partial1 — Weak0 — Missing
Goal, planning, and action choiceGoal stays explicit; subgoals are observable/revisable; selected actions match current need, available evidence, and least-authority constraintsCorrect goal/subgoal/action flow with minor gapsFunctional flow but planning or action rationale is weakFrequent irrelevant/drifting actionsNo coherent goal-directed loop
State and memoryWorking state is typed/selective with provenance; control fields are protected; persistent writes and retrieval follow scope/freshness/privacy policyCorrect state/memory boundaries with minor gapsState works but provenance or memory policy is shallowMostly transcript-driven or stale memory handlingNo reliable state/memory boundary
Authority, approval, and boundednessPermissions are external to model output; approvals bind exact actions and expire; stop rules cover success/blocker/denial/repetition/budget with zero unsafe continuationCorrect boundaries with minor gapsMost controls exist but one important edge is weakPrompt wording carries major control responsibilityUnauthorized action, approval bypass, or unbounded loop can pass
Failure analysis and evaluationFixed trajectories cover key failure classes; earliest-boundary diagnosis is explicit; metrics separate task, selection, safety, memory, drift, and efficiency; release is recomputedStrong evaluation with small omissionsFixed cases exist but slices/diagnosis or recomputation is limitedMostly anecdotal demos or final-answer-only checksNo meaningful trajectory evaluation
Reproducibility and communicationStandard-library deterministic path, versioned policy/evaluator/fixtures, exact commands, debug record, limitations, and optional live-path differences are clearReproducible with minor omissionsSeveral identities or limitations missingManual screenshots/output onlyNo reproducible package

Objective checks

The validator checks active/pending subgoal choice, permission-filtered action selection, protected working-memory updates, memory-write policy, exact/expiring approval binding, stopping rules, explicit transitions, raw-trajectory metric recomputation, goal-drift detection, repeated-action bounds, prohibited memory-write rejection, unauthorized execution rejection, and release-rule recomputation.

The intentional failure fixture keeps the task outcomes correct while inserting control failures. A passing implementation must reject it; final-answer success cannot compensate for authority, memory, or boundedness violations.

← Back to all projects