Agent Loops
Goal
Implement the outer control loop for an agent with an explicit step budget, terminal states, and progress checks so repeated decision making cannot run without a bound.
The observe–decide–act cycle describes one turn. An agent loop repeats that cycle until a stopping condition is reached.
A minimal sketch looks like this:
while not done:
observation = observe(state)
decision = decide(state, observation)
result = act(decision)
state = update(state, result)
The sketch is useful, but it is incomplete. If done never becomes true, the run can continue forever. If the same failed action is chosen repeatedly, cost grows while progress stays at zero.
Agent loop: observe, decide, act, repeat—or stop
The model proposes the next step, but application code decides what is allowed, when to continue, and when the run must stop.
| Terminal reason | Meaning |
|---|---|
SUCCESS | The goal is satisfied. |
BLOCKED | Required information or permission is unavailable. |
DENIED | Policy or human approval rejected the action. |
FAILED | A non-recoverable error ended the run. |
BUDGET_EXCEEDED | The configured step or cost limit was reached. |
Put the bound outside the model
A prompt can say “finish within eight steps,” but application code should enforce the number.
For example:
for step in range(max_steps):
...
else:
stop_reason = "max_steps"
The model may suggest continuing on step nine. The loop controller should still stop if the configured budget is eight.
This is the same principle as tool permissions: important safety and cost boundaries should not depend only on generated text.
Terminal states should be explicit
A run should not merely “end.” Record why it ended.
Useful terminal states include:
SUCCESS— the goal is satisfied;BLOCKED— required information or permission is unavailable;DENIED— policy or human approval rejected the next action;FAILED— a non-recoverable error occurred;BUDGET_EXCEEDED— the step or cost limit was reached.
Different stop reasons need different follow-up behavior. A blocked run might ask the user for missing information. A denied action should not be retried as if it were a timeout.
Detect repeated actions
Imagine an agent that calls search_orders("4172") five times and receives the same result each time. The step budget eventually stops it, but repetition detection can stop it earlier.
One simple rule is:
if same normalized action appears 3 times without new evidence:
stop_reason = repeated_action
The threshold is application-specific. The key idea is to make “no progress” observable.
Count attempts even when tools fail
If a timeout occurs, the attempt still consumed time and possibly money. Step budgets should not reset just because an action did not return a normal result.
Retry rules from Level 10 remain separate. A retry may be allowed, but it still belongs inside the outer agent budget.
The loop controller owns continuation
The model can propose continue, answer, ask_user, or a tool action. The controller maps those proposals into legal transitions.
This means a model output cannot invent a new terminal state or skip a required approval. It chooses within a controlled vocabulary.
Predict
Run the local Lab
Run:
python labs/notebooks/level-11/l11-03-agent-loop.py
The script contains one run that succeeds and one run that repeats an action.
- Run it unchanged and record the terminal reason for both traces.
- In
run(...), the default ismax_steps=4. Before editing, predict which trace will change terminal reason if only the step budget becomes 2. - Change only
max_steps=4tomax_steps=2, then rerun. - Compare the terminal reasons. Explain why the same proposed actions can end differently when the controller's outer budget changes.
Loading lab…
Write the core logic yourself
Open:
labs/notebooks/level-11/l11-03-agent-loop-exercise.py
Implement success, repetition, and step-budget stopping in the outer loop. The starter deliberately contains no completed loop logic.
Run the starter after each change:
python3 labs/notebooks/level-11/l11-03-agent-loop-exercise.py
A correct implementation ends with a PASS: marker. Only after you have a working version, compare your approach with the solved deterministic script used by the Level smoke tests.
Optional: put a real model inside the bounded loop
Run:
python labs/real-model/l11_agent_loop_real_model.py --max-steps 3
Here the model can choose the order-status tool and then see the tool result, but the controller remains in charge of continuation. It rejects undeclared tools, rejects malformed arguments, blocks an identical repeated action, and stops when the configured step budget is exhausted.
Rerun with --max-steps 1. The model has not changed, but the outer controller now permits less work. That contrast is the point: model reasoning happens inside application-owned bounds.
Quick Check
Explain it back
Describe two different ways the same agent can stop: one because the task is complete and one because useful progress has stopped. Explain which application field would distinguish those outcomes.
Key Takeaways
- Agent loops need application-enforced bounds.
- Terminal states should record why a run ended.
- Repetition detection can identify stalled progress before the full budget is used.
- Failed attempts still consume the outer run budget.
- The loop controller decides which continuation and stop transitions are legal.
Next Lesson
Next, learn how a generated plan can organize a longer goal without becoming an unquestioned script.
References
Completion is stored locally on this device.