Skip to main content

Mini Checkpoint — Build the Bounded Loop

You have now separated the main parts of an agent run:

goal
→ observation
→ decision
→ permitted action
→ new observation
→ explicit state update
→ continue or stop

Before adding more sophisticated tool selection and critique, check that these foundations are solid.

1. Classify the system​

For each system, write single call, fixed workflow, or agent-like loop.

  1. A model writes one answer and stops.
  2. A script always runs search, then rerank, then summarize.
  3. A system searches, examines the result, then decides whether to search again, ask the user, or stop.

Explain what evidence makes item 3 different from item 2.

2. Trace two cycles​

Use this starting state:

goal = resolve order 4172
order_status = unknown
refund_eligible = unknown
step_count = 0

Write two cycles with these headings:

observation:
decision:
action:
state update:

Your trace must keep a proposed action separate from the tool result.

3. Add two bounds​

Choose a maximum step count and a repeated-action rule. Explain which one would stop this run first:

step 1: search order 4172 → same result
step 2: search order 4172 → same result
step 3: search order 4172 → same result
step 4: search order 4172 → same result

A good answer names the exact terminal reason.

4. Revise a plan​

Start with:

1. get order status
2. check refund eligibility
3. request approval
4. issue refund

Now add the observation order_status = shipping. Mark each subgoal as pending, completed, blocked, or skipped and explain which plan change is supported by the new evidence.

5. Separate working and long-term memory​

Place each item in working state, long-term candidate, audit history only, or reject:

  • current step count;
  • latest trusted order status;
  • raw tool response from step 1;
  • user's stable contact preference from a trusted profile;
  • a model-generated guess about the user's bank account;
  • approval token bound to the current refund action.

Explain one item whose category could change depending on privacy or retention policy.

Check your reasoning after you try

A strong trace makes the control boundary visible:

  • Item 1 is a single call, item 2 is a fixed workflow, and item 3 is an agent-like loop because a later action depends on a new observation.
  • In a two-cycle trace, keep decision separate from action result. The model can propose a search; the tool result becomes the next observation only after execution.
  • A repeated-action rule can stop the sample search loop before a larger step budget. Record an exact terminal reason such as repeated_action_limit.
  • If order_status = shipping, the plan should change only where that evidence matters. Do not continue refund steps just because they were in the original plan.
  • Working state can include the current step count and trusted current status. Stable preferences may be long-term candidates under policy. Raw tool output can remain audit evidence. An unsupported bank-account guess should be rejected.

Applied check: suppose the third search returns a different order status. A bounded loop should update state and reconsider the plan; it should not stop merely because the previous two searches matched.

Success check​

You are ready to continue if you can explain all five ideas without relying on the word “autonomous”:

  • later actions can depend on new observations;
  • the application owns loop bounds;
  • plans are revisable;
  • working state is selective;
  • persistent memory is governed evidence, not automatic truth.

The next lessons use these foundations to choose tools, critique decisions, stop safely, and place humans at important control points.

Lesson actions

Completion is stored locally on this device.

View progress