Skip to main content
L11.7

Tool Selection

Goal

Choose the next tool from the current subgoal, trusted state, tool capabilities, and policy constraints instead of selecting by name similarity alone.

Level 10 taught you to expose narrow tools and validate their arguments. An agent adds a new question: which tool should be considered next?

A model may see several tool descriptions at once. Some tools can read information, some can change external state, and some may not be permitted for the current principal. Good selection therefore depends on more than whether a tool name sounds related to the user's words.

Start from the current subgoal​

Suppose the current subgoal is:

determine whether order 4172 is eligible for refund

Available tools are:

get_order(order_id)
get_refund_policy(region, item_type)
refund_order(order_id, amount)
send_message(channel, text)

The side-effect tool refund_order is not the first choice. The current subgoal needs evidence, so the two read-only tools are stronger candidates.

A useful selector asks, “What missing evidence would let this subgoal move forward?”

Capability is not permission​

A tool can be technically capable of performing an action and still be unavailable to the current run.

If refund_order exists but the current role cannot use it, the selection layer should not pretend that permission is merely a later detail. The available-tool set given to the model can already be filtered by policy.

Defense in depth is still useful: the execution boundary should check permission again before running a side effect.

Prefer the narrowest tool that answers the question​

If the task needs one order status, a tool that returns one order is usually easier to reason about than a broad database-query tool.

The narrow choice reduces argument ambiguity, limits result size, and makes success easier to evaluate. It also reduces the number of unintended operations the model could ask for.

This is the same least-authority idea from Level 10, now applied during repeated next-action selection.

Record alternatives when selection is uncertain​

Sometimes two tools look plausible. Rather than hiding that uncertainty, record enough information to inspect it.

For example:

candidate: get_order
supports: order status, region, item type
candidate: get_refund_policy
supports: policy once region/item type are known
chosen: get_order
reason: required policy inputs are still unknown

That trace reveals why one tool was selected before the other.

Tool results can change the candidate set​

After get_order returns the region and item type, get_refund_policy becomes useful. After the policy says approval is required, the next action may be request_approval rather than another information tool.

Tool selection is therefore state-dependent. The correct tool at step 1 may be wrong at step 3.

Do not reward tool use for its own sake​

If current state already contains a fresh trusted answer, another lookup may be unnecessary. A selector should be allowed to choose “answer” or “stop” rather than forcing a tool call.

More tool calls can increase cost, latency, and failure surface without adding evidence.

Predict

The current subgoal is to learn an order's region and item type before checking policy. Which tool is the strongest first choice?

Run the local Lab​

Run:

python labs/notebooks/level-11/l11-07-tool-selection.py

The script scores each tool: +10 for every field it provides that is still missing, −5 if it has side effects, and −100 if the tool is not permitted.

  1. Run it unchanged. Read still missing: ['item_type', 'refund_policy', 'region'].
  2. Read the ranking: get_order scores 20 (it fills region and item_type), get_refund_policy scores 10, and refund_order scores -100 because it is not permitted. selected: get_order.
  3. Now pretend the order lookup already happened. Change known_fields = {"order_id"} to known_fields = {"order_id", "region", "item_type"} and rerun.
  4. Now still missing: is ['refund_policy']. get_order drops to 0 because it would only repeat known facts, and selected: becomes get_refund_policy.

The same tool list gives a different best next step once the state changes. Good selection depends on what is still unknown.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-11/l11-07-tool-selection-exercise.py

Write the scoring rule so useful permitted read tools beat irrelevant or forbidden capabilities. A forbidden side-effecting tool must never win simply because it could produce a desired field.

Run the starter after each change:

python3 labs/notebooks/level-11/l11-07-tool-selection-exercise.py

A correct implementation ends with a PASS: marker. Only after you have a working version, compare your approach with the solved deterministic script used by the Level smoke tests.

Quick Check

1. What should primarily drive an agent's next-tool choice?
2. Why can the best tool change later in the same run?
3. When should an agent be allowed to choose no tool?

0 of 3 questions answered.

Explain it back​

Choose four tools from an application you know. For one subgoal, rank the tools and explain which missing fact or permission makes the top choice stronger than the others. Then change one state fact and explain how the ranking changes.

Key Takeaways

  • Tool selection begins from the current subgoal and missing evidence.
  • Capability and permission are different; filter and recheck both.
  • Narrow tools usually make selection and evaluation easier.
  • Record why a candidate was chosen when several tools are plausible.
  • “No tool” can be the correct choice when current state is sufficient.

Next Lesson

Next, add a focused critique step that checks a proposed decision against evidence and rules without creating an endless second loop.

References

Lesson actions

Completion is stored locally on this device.

View progress