Skip to main content
L15.7

Tool and Agent Security

Goal

Apply least privilege, trusted identity, policy checks, approval, and audit evidence to tool-using and agentic workflows.

Keep identity, proposal, and policy separate​

Picture a student handing a request slip to an office worker: "Please unlock room 12." The handwriting on the slip can describe the requested action, but it cannot declare, "I am the principal, so I am allowed." The office checks the student's real identity and the school rule before touching the lock.

Tool and agent security uses the same separation. The principal is the trusted identity on whose behalf an action would occur. The model's proposal describes a requested action. The policy is trusted application logic that decides whether that principal may perform that action on that resource. Keeping those three roles separate prevents generated text from granting itself authority.

A tool changes the stakes of a model output. Without a tool, a bad answer may mislead a user. With a tool, the same model can propose sending a message, changing a record, running code, or transferring data. The application must decide whether that proposal is allowed.

{"principal":"user-42","proposal":{"tool":"refund","order_id":"4172","amount":20},"policy_decision":"deny"}

Never let generated text grant authority​

Do not let the proposal redefine the principal. A model should not be able to add "role":"admin" or switch tenant_id and thereby gain authority. Trusted identity comes from authenticated application state. Model-generated arguments may select task details only within the scope the principal already has.

Use least privilege and approval at real boundaries​

Least privilege means exposing only the tools and permissions needed for the current task. A support assistant that only reads order status should not receive a generic network client or shell. A calendar helper that can create events may still need separate approval before inviting external people.

Approval is most useful at irreversible or high-impact boundaries. Do not ask for confirmation after the action already happened. Present the user with the action, destination, and important parameters before execution. The confirmation itself should bind to that exact proposal so a later model turn cannot silently change the amount or recipient.

Bound loops and delegated authority​

Agent loops add persistence and repetition. A single unsafe proposal can be rejected, but an unconstrained agent may retry with modified arguments. Stopping conditions, retry budgets, and immutable principal information prevent “helpful” repair loops from becoming privilege escalation.

Delegation adds another identity question. If one agent asks another to perform work, the receiving side needs to know which authority is being delegated and which is not. “Another agent asked me” is not sufficient authorization. Preserve task identity, principal, allowed scope, and evidence of the delegation boundary.

Layer security principles with local architecture​

NIST SP 800-207's zero-trust model is useful as a general security principle: do not grant broad trust merely because a component is inside a network boundary. For AI systems, apply that mindset to model outputs and agent messages too. Each protected resource request should be evaluated using trusted identity and policy.

OWASP and MITRE guidance can help enumerate tool and agent failure modes, but architecture still determines the control. A model prompt cannot reliably replace a permission check. A tool schema cannot replace resource authorization. A sandbox cannot replace a limit on which data may enter or leave it.

Record and test denied paths​

Audit evidence should explain the decision. Record principal, requested action, resource, policy version, approval state, execution result, and correlation ID. Avoid logging secrets or complete sensitive payloads when a classification or digest is enough.

Security tests should cover denied paths, not only successful ones. Verify that an unauthorized principal is rejected, that changing model-generated role fields does not change authority, that approval is required where declared, and that retries cannot bypass the same policy.

Predict

A model retries a denied tool call and changes the generated role from user to admin. What should the application do?

Run the local Lab​

Run:

python3 labs/notebooks/level-15/l15-07-tool-agent-security.py

The Lab evaluates a generated tool proposal against a trusted principal and application policy.

  1. Run it unchanged. Record trusted_role: support, the proposal's generated role: admin, and authorized: true for the allowed read_order action.
  2. Before editing, predict whether authorization should change if only the model-generated role string changes.
  3. Change only proposal["role"] from "admin" to "owner", then rerun.
  4. Confirm generated_role_ignored changes but trusted_role and authorized do not. Explain why authority comes from the trusted principal, not from model-generated text.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-15/l15-07-tool-agent-security-exercise.py

Implement authorization from the trusted principal role and application policy. The generated proposal includes a spoofed admin role on purpose; your code must ignore it.

Run:

python3 labs/notebooks/level-15/l15-07-tool-agent-security-exercise.py

The starter intentionally stops at TODO until you implement the missing logic. A correct solution reaches the final PASS: marker. Use the solved deterministic Lab as a comparison only after your own attempt.

Quick Check

1. Which value should determine a tool caller's authority?
2. Why bind human approval to the exact proposed action?
3. What should a security regression suite include for agent tools?

0 of 3 questions answered.

Explain it back​

Describe a tool call as principal + proposal + policy. Then explain what happens when the model changes role, tenant, or approval fields in the proposal.

Key Takeaways

  • Model output proposes actions; trusted application policy authorizes them.
  • Principal identity must not be writable by the model.
  • Use least privilege and approvals at meaningful side-effect boundaries.
  • Agent retries and delegation must preserve the same security constraints.
  • Audit evidence should make authorization decisions reviewable without leaking secrets.

Next Lesson

Next, L15.8 — Privacy, Data Governance, and Retention follows data through collection, use, access, storage, and deletion.

References

Lesson actions

Completion is stored locally on this device.

View progress