Skip to main content
L10.5

Error Handling and Retries

Goal

Classify tool failures, retry only when the failure is plausibly temporary or repairable, and enforce bounded retry and idempotency rules.

Tool workflows fail for ordinary software reasons. A service may time out. A record may not exist. Arguments may be invalid. A permission check may reject the request.

The important skill is not “retry on error.” It is deciding which error changes if you try again.

Separate failure categories​

A useful reduced set is:

validation → arguments are wrong
permission → action is not allowed
not_found → requested resource is absent
rate_limit → service asks you to slow down
timeout → result is unknown/late
server → service failed internally

Blindly retrying validation or permission failures usually repeats the same mistake. A timeout or rate limit may be retryable under a bounded policy.

Repair is different from retry​

If an argument is missing, the next step may be to construct a corrected request or ask the user. Sending the identical invalid call again is not progress.

Record whether the next attempt changed anything meaningful.

Side effects need idempotency thinking​

Suppose a payment tool times out after the request is sent. Did the payment happen? If you retry without an idempotency mechanism, you may create a duplicate charge.

Read-only requests are often easier to retry than actions. For side effects, the workflow needs an operation ID, idempotency key, status check, or another application-specific protection.

Set a retry budget​

Every retry loop should have a stopping condition such as:

maximum attempts = 3
maximum elapsed time = 10 seconds
retryable categories = {timeout, rate_limit}

After the budget is exhausted, return a clear failure or escalate. Do not let the model decide to loop forever.

Backoff reduces synchronized pressure​

When a service asks clients to slow down, retrying instantly can make the problem worse. Increasing delays—often with jitter—can reduce contention.

The exact timing policy is an application/runtime detail. The stable principle is: retry behavior should be explicit, bounded, and observable.

Walk through one timeout carefully​

Suppose get_order times out before returning data. Because it is read-only, a bounded retry is usually straightforward: no external state is changed by repeating the lookup. Now suppose refund_order times out after the request left your process. The missing response does not tell you whether the refund failed or succeeded. Repeating it immediately can duplicate the side effect.

A safer workflow first checks whether the operation has an idempotency key or a way to query completion. If neither exists, the correct next step may be escalation instead of automatic retry. The same error label—timeout—therefore leads to different actions because the operation semantics differ.

Track attempt identity​

Record attempt number, call ID, operation/idempotency ID, error category, and next decision. This makes it possible to distinguish three independent failures: the service failed repeatedly, the workflow retried when it should not have, or the workflow exceeded its declared retry budget.

Predict

A cancel_order call times out after it may have reached the service. What should the workflow consider before repeating the call?

Run the local Lab​

python labs/notebooks/level-10/l10-05-retry-policy.py

Each case is (category, attempt, side_effect, idempotent), and the loop allows max_attempts = 2.

  1. Run the command. You should see validation -> stop_or_repair, timeout -> retry, timeout -> check_completion_before_retry (a timeout on a side effect that is not idempotent), and rate_limit -> stop_budget (already at attempt 2).
  2. Use up the budget for the first timeout: change ("timeout", 0, False, False), to ("timeout", 2, False, False), and rerun. That line becomes timeout -> stop_budget: the same kind of failure stops once the attempts are used up.
  3. Change it back. Now change the first case, ("validation", 0, False, False),, to ("timeout", 0, False, False), and rerun. It becomes retry.
  4. Explain why: a validation error will fail the same way every time, so retrying cannot help, while a timeout may succeed on a second try. Change the case back afterward.

Loading lab…

Quick Check

1. Which failure is usually fixed by repeating the identical request?
2. Why are side-effect retries special?
3. What belongs in a retry policy?

0 of 3 questions answered.

Key Takeaways

  • Classify failures before deciding to retry.
  • Repairing an invalid request is different from repeating it.
  • Side-effect retries require idempotency or completion checks.
  • Retry budgets need explicit stopping conditions.
  • Backoff and observability are part of reliable retry behavior.

Next Lesson

Next, represent tool workflows as explicit states and allowed transitions so retries, approvals, and stopping become easier to reason about.

References

Lesson actions

Completion is stored locally on this device.

View progress