본문으로 건너뛰기
L12.3

Retries, Idempotency, and Recovery

Goal

Design retry rules that distinguish transient failures from unsafe repetition and use idempotency keys to prevent duplicate side effects during recovery.

The dangerous part of a timeout​

Think about pressing a "pay $20" button and seeing a timeout. You do not know whether the payment failed before reaching the server or succeeded just before the reply was lost. Pressing the button again blindly could pay twice.

That uncertainty is why retries and duplicate safety must be designed together. An idempotency key gives repeated attempts one stable operation identity: "this is still the same $20 payment," not "please create another payment." The word is technical, but the goal is simple—when the same intended operation is retried, the system should be able to recognize the duplicate and avoid an unintended second side effect.

The next question is which failures are worth retrying at all.

Give retries one operation identity​

Idempotency does not mean every action naturally has no effect when repeated. It means the system gives repeated attempts a stable identity and defines how duplicates are handled. Read-only lookups are often safe to repeat. Charges, refunds, sends, deletes, and writes need stronger care because repetition may create new external state.

operation_id = pay:order-4172:20
attempt 1 → timeout after send
attempt 2 → same operation_id → confirm/reuse prior effect

Classify failures before retrying​

Retry policy should classify failures. A timeout may be retryable. Invalid arguments should usually be fixed rather than retried unchanged. A permission denial should stop, not back off and try again. A human rejection should remain rejected. Treating all errors as retryable turns a bounded system into a persistence machine that ignores the meaning of failure.

Backoff controls when retries happen by waiting longer after repeated failures. A common pattern also adds jitter, a small timing variation, so many workers do not retry at the same instant. This lesson uses deterministic intervals for learning, but the important concept is that retry timing is part of controller policy, not something the model improvises.

Recover from evidence, not guesses​

Recovery after a crash should inspect durable evidence before deciding to repeat an operation. Useful evidence includes the operation ID, request arguments, last known state, external receipt, and whether the operation can be queried by identity. The safest next step may be reconciliation: check the external system or another durable record to learn what happened instead of performing the effect again.

Exactly-once execution across distributed systems is difficult to guarantee end to end. A more practical design often combines at-least-once task delivery—the same task may be delivered more than once—with idempotent effects and durable completion records. That way a task may be delivered twice while the external side effect still happens once for the same operation identity.

Agent reasoning should not be responsible for deciding whether an uncertain side effect occurred. The model can summarize observations, but the controller should use operation records and tool-specific reconciliation logic. This keeps recovery grounded in external facts.

A good test for a retry design is to inject a crash at awkward moments: before the request, after sending but before recording the response, and after recording completion. The correct behavior should be explainable at each point. If one crash position can create a duplicate effect, the recovery boundary still needs work.

Predict

A payment request times out after being sent. What is the safest next controller step when the service supports lookup by operation ID?

Run the local Lab​

Run:

python labs/notebooks/level-12/l12-03-retry-idempotency.py

The script simulates retrying the same refund through an idempotency ledger.

  1. Run it unchanged with reuse_operation_id=True. Confirm both attempts return the same receipt and effect_count: 1.
  2. Before editing, predict the receipt IDs and effect count if the retry gets a new operation ID.
  3. Change only reuse_operation_id=True to reuse_operation_id=False, then rerun.
  4. Confirm the ledger creates a second receipt/effect and the final safety assertion fails. Restore the starter afterward. The retry payload stayed the same; only operation identity changed.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-12/l12-03-retry-idempotency-exercise.py

Implement stable operation identity so the same retry returns the original receipt instead of creating a second side effect.

Run:

python3 labs/notebooks/level-12/l12-03-retry-idempotency-exercise.py

The starter intentionally fails at its TODO boundary. A completed implementation ends with a PASS: marker. Compare with the solved deterministic Lab only after attempting the implementation yourself.

Quick Check

1. What problem does an idempotency key solve?
2. Which failure should normally stop rather than retry unchanged?
3. Why can at-least-once task delivery still be useful?

0 of 3 questions answered.

Explain it back​

Describe a refund operation that times out after sending. Explain the durable fields you would inspect and how the operation ID changes the recovery decision.

Key Takeaways

  • Retry only failures that are meaningful to retry.
  • Use stable operation identities for side effects.
  • Reconcile uncertain effects before blindly repeating them.
  • Permission denial and human rejection are not transient errors.
  • Crash injection is a powerful way to test recovery logic.

Next Lesson

Next, L12.4 — Tracing Agent Decisions makes these retries, transitions, and effects visible as structured evidence.

References

Lesson actions

Completion is stored locally on this device.

View progress