Reflection and Critique
Goal
Use a bounded critique step to check a proposed action against explicit evidence and rules, and distinguish useful correction from unbounded self-review.
An agent can make a plausible decision for a bad reason. One way to catch some mistakes is to insert a critique step before an important action.
A critique is not a magical truth detector. It is another decision process. It can also be wrong. Its value comes from asking a narrower question with explicit evidence and criteria.
Critique the decision, not the agent's personality
Weak critique prompt:
“Think harder. Are you sure?”
Stronger critique task:
proposed action: refund_order(order_id=4172, amount=80)
check:
- is refund eligibility present in trusted state?
- is the principal permitted to request this tool?
- is exact approval present for order 4172 and amount 80?
- would this action exceed the remaining step budget?
return: pass or named issue
The stronger form turns critique into a check against inspectable conditions.
Keep deterministic checks deterministic
If the approval amount is wrong, application code can compare values exactly. Do not ask a language model to “reflect” on whether two dictionaries are equal.
Use critique for judgment-heavy questions such as whether the proposed action addresses the current subgoal or whether the evidence cited by the model actually supports its interpretation.
Then keep schema, permission, approval binding, counters, and legal transitions in deterministic code.
Critique should have a budget
A common failure mode is:
draft
→ critique
→ revise
→ critique
→ revise
→ critique
→ ...
The system appears careful but may never act or stop.
A bounded policy might allow one critique before a high-impact action and at most one revision. If the revised proposal still fails deterministic checks, stop or escalate.
The exact number depends on the task. The stable principle is that self-review needs the same kind of budget as the outer agent loop.
Record critique outcomes
Useful outcomes include:
pass;missing_evidence;goal_mismatch;unsupported_claim;needs_human_review.
These labels make critique measurable. “The reflection was good” is difficult to test.
You can later ask whether critique caught real errors or merely added latency.
Do not let critique override policy
A critique model may say an unauthorized action is reasonable. That does not grant permission.
The critique can recommend “request approval” or “stop,” but trusted policy still controls action authority.
Likewise, a critique should not rewrite the user's goal to make the current plan look successful.
Predict
Run the local Lab
Run:
python labs/notebooks/level-11/l11-08-critique.py
The script checks three proposals against two rules: the proposal must include evidence, and it must match the active subgoal. The controller then decides what to do with each result.
- Run it unchanged. Read:
p1 issues: ['needs_evidence'] => revise (1 of 1)p2 issues: ['matches_subgoal'] => revise (1 of 1)p3 issues: none => accept
- Open the script and change
max_revisions = 1tomax_revisions = 0. Rerun. p1andp2still show the same issues, but the action becomesrecord issue and escalate (no revision allowed).p3is still accepted.
The critique still finds the problem. The revision budget only controls whether the agent may try again or must hand the issue to someone else.
Loading lab…
Quick Check
Explain it back
Take one important action in your project. Write a three-item critique rubric for the judgment-heavy parts, then list two deterministic checks that should happen separately. State the maximum number of critique/revision cycles you would permit.
Key Takeaways
- Critique is another fallible decision step, not a truth oracle.
- Focus critique on explicit evidence and criteria.
- Keep exact policy, schema, counter, and transition checks deterministic.
- Bound critique and revision so self-review cannot loop forever.
- Measure critique with named outcomes and correction evidence.
Next Lesson
Next, make stopping a first-class decision with success, blocker, repetition, denial, and budget conditions.
References
Completion is stored locally on this device.