Tool Selection
Goal
Choose the next tool from the current subgoal, trusted state, tool capabilities, and policy constraints instead of selecting by name similarity alone.
Level 10 taught you to expose narrow tools and validate their arguments. An agent adds a new question: which tool should be considered next?
A model may see several tool descriptions at once. Some tools can read information, some can change external state, and some may not be permitted for the current principal. Good selection therefore depends on more than whether a tool name sounds related to the user's words.
Start from the current subgoal
Suppose the current subgoal is:
determine whether order 4172 is eligible for refund
Available tools are:
get_order(order_id)
get_refund_policy(region, item_type)
refund_order(order_id, amount)
send_message(channel, text)
The side-effect tool refund_order is not the first choice. The current subgoal needs evidence, so the two read-only tools are stronger candidates.
A useful selector asks, “What missing evidence would let this subgoal move forward?”
Capability is not permission
A tool can be technically capable of performing an action and still be unavailable to the current run.
If refund_order exists but the current role cannot use it, the selection layer should not pretend that permission is merely a later detail. The available-tool set given to the model can already be filtered by policy.
Defense in depth is still useful: the execution boundary should check permission again before running a side effect.
Prefer the narrowest tool that answers the question
If the task needs one order status, a tool that returns one order is usually easier to reason about than a broad database-query tool.
The narrow choice reduces argument ambiguity, limits result size, and makes success easier to evaluate. It also reduces the number of unintended operations the model could ask for.
This is the same least-authority idea from Level 10, now applied during repeated next-action selection.
Record alternatives when selection is uncertain
Sometimes two tools look plausible. Rather than hiding that uncertainty, record enough information to inspect it.
For example:
candidate: get_order
supports: order status, region, item type
candidate: get_refund_policy
supports: policy once region/item type are known
chosen: get_order
reason: required policy inputs are still unknown
That trace reveals why one tool was selected before the other.
Tool results can change the candidate set
After get_order returns the region and item type, get_refund_policy becomes useful. After the policy says approval is required, the next action may be request_approval rather than another information tool.
Tool selection is therefore state-dependent. The correct tool at step 1 may be wrong at step 3.
Do not reward tool use for its own sake
If current state already contains a fresh trusted answer, another lookup may be unnecessary. A selector should be allowed to choose “answer” or “stop” rather than forcing a tool call.
More tool calls can increase cost, latency, and failure surface without adding evidence.
Predict
Run the local Lab
Run:
python labs/notebooks/level-11/l11-07-tool-selection.py
The script scores each tool: +10 for every field it provides that is still missing, −5 if it has side effects, and −100 if the tool is not permitted.
- Run it unchanged. Read
still missing: ['item_type', 'refund_policy', 'region']. - Read the ranking:
get_orderscores20(it fillsregionanditem_type),get_refund_policyscores10, andrefund_orderscores-100because it is not permitted.selected: get_order. - Now pretend the order lookup already happened. Change
known_fields = {"order_id"}toknown_fields = {"order_id", "region", "item_type"}and rerun. - Now
still missing:is['refund_policy'].get_orderdrops to0because it would only repeat known facts, andselected:becomesget_refund_policy.
The same tool list gives a different best next step once the state changes. Good selection depends on what is still unknown.
Loading lab…
Write the core logic yourself
Open:
labs/notebooks/level-11/l11-07-tool-selection-exercise.py
Write the scoring rule so useful permitted read tools beat irrelevant or forbidden capabilities. A forbidden side-effecting tool must never win simply because it could produce a desired field.
Run the starter after each change:
python3 labs/notebooks/level-11/l11-07-tool-selection-exercise.py
A correct implementation ends with a PASS: marker. Only after you have a working version, compare your approach with the solved deterministic script used by the Level smoke tests.
Quick Check
Explain it back
Choose four tools from an application you know. For one subgoal, rank the tools and explain which missing fact or permission makes the top choice stronger than the others. Then change one state fact and explain how the ranking changes.
Key Takeaways
- Tool selection begins from the current subgoal and missing evidence.
- Capability and permission are different; filter and recheck both.
- Narrow tools usually make selection and evaluation easier.
- Record why a candidate was chosen when several tools are plausible.
- “No tool” can be the correct choice when current state is sufficient.
Next Lesson
Next, add a focused critique step that checks a proposed decision against evidence and rules without creating an endless second loop.
References
- Schick et al., Toolformer.
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models.
Completion is stored locally on this device.