Observe, Decide, Act
Goal
Trace one agent cycle by separating observations, decisions, actions, and resulting state updates instead of treating the run as one opaque model conversation.
An agent loop becomes easier to reason about when each step has a job. The simplest useful cycle is observe → decide → act.
An observation is evidence available at the current step: a user goal, a tool result, an approval response, a timeout, or a value already stored in state. A decision chooses what should happen next. An action is the permitted operation that actually runs.
After the action, a new observation becomes available and the cycle can begin again.
Follow one cycle with concrete data
Start with this state:
goal = "resolve order 4172"
order_status = unknown
approval = none
step = 0
The system observes that the order status is unknown. The decision is get_order. Application code validates and executes the tool. The tool returns {"status":"delivered","damaged":true}.
That result is not merely text to append somewhere. It is a new observation that should update state in a controlled form:
order_status = delivered
damaged = true
step = 1
Now the next decision has different evidence than the first decision.
Do not mix observation and action
A common design mistake is to store something like “refund the order” as if it were an observation. That phrase is a proposed action, not external evidence.
Keeping categories separate helps answer debugging questions. Did the tool return the wrong fact? Did the model choose the wrong action from a correct fact? Did application code execute an action that policy should have denied?
If the trace blends everything into one transcript, those questions become much harder.
State is the bridge between cycles
The complete conversation history can be useful, but an agent should not depend on an ever-growing text transcript as its only state representation. Important facts should have explicit fields when practical.
For the running example, useful state might include:
- goal;
- current subgoal;
- trusted order facts;
- requested and completed actions;
- approval status;
- step count;
- terminal status.
This does not eliminate model context. It gives the application a stable record that does not depend on the model correctly re-reading every old sentence.
Decisions should name evidence
When possible, record why an action was chosen using inspectable evidence.
For example:
decision: check_refund_policy
because:
order_status = delivered
damaged = true
policy_status = unknown
This is more useful than a vague note such as “the agent thought it was appropriate.” You can test whether the required evidence was present.
Actions still cross the Level 10 boundary
The decision refund_order is not the refund. Before an external side effect, application code still checks the tool schema, principal permissions, required approval, and legal workflow transition.
The observe–decide–act loop extends the number of decision points. It does not remove the action boundary.
Predict
Run the local Lab
Run:
python labs/notebooks/level-11/l11-02-observe-decide-act.py
The script starts with the action get_order. After each action it records the observation, updates state, and then decides the next action from that observation.
- Run it. You should see two cycles:
get_order→ decisioncheck_refund_policy, thencheck_refund_policy→ decisionrequest_approval. The last line isstopped at: request_approval after 2 cycle(s). - In
TOOL_RESULTS, change theget_orderresult from"status": "delivered"to"status": "shipping"and rerun. - Now the first decision is
report_shipping, and the run stops after 1 cycle. The policy check never happens, because the new observation made it unnecessary.statestill recordsorder_statusandstepexplicitly. - Change the status back to
delivered.
Loading lab…
Quick Check
Explain it back
For one two-step task, write four lines for each cycle: observation, decision, action, state update. If you cannot tell which line contains evidence and which contains a proposal, separate them more clearly.
Key Takeaways
- Observe, decide, and act are different roles in the loop.
- New tool results become observations that update controlled state.
- Explicit state fields make important facts and counters inspectable.
- Decision records are stronger when they name the evidence used.
- Side effects still require application-level validation and authority checks.
Next Lesson
Next, wrap repeated cycles in an explicit agent loop and make the loop stop safely when progress ends or a budget is reached.
References
Completion is stored locally on this device.