본문으로 건너뛰기

Checkpoint — Durable Agent Foundations

Before moving into parallel work and delegation, make sure the first half of Level 12 forms one coherent picture.

A bounded agent from Level 11 decides what to do next. The Level 12 harness makes that work survive failures and remain inspectable. You should now be able to explain why temporary process memory is not enough, how a queue separates task lifetime from worker lifetime, and why durable state must include task identity, state, completed operation IDs, and other facts required for safe resume.

You should also be able to explain a timeout after a side effect. The correct question is not simply whether to retry. First determine whether the original operation has a stable identity and whether completion can be reconciled from durable or external evidence. A retry without that evidence can duplicate an effect.

Tracing should give you observable application evidence: task IDs, spans, action names, policy decisions, attempts, timings, and terminal reasons. It should not require storing hidden reasoning. You should know why sensitive arguments may need redaction and why version information helps connect behavior changes to releases.

Finally, separate three security ideas. Permission policy decides whether the requested action is allowed. A sandbox limits what the executing process can reach. Memory policy controls which stored information may influence later decisions. These controls reinforce one another but should not be collapsed into one boolean.

Self-check​

You are ready to continue if you can answer these without looking back:

  1. What information must survive a worker crash?
  2. Why can a queue redeliver work without making duplicate side effects acceptable?
  3. How does a stable operation ID change retry behavior?
  4. Which trace fields help distinguish a policy denial from a tool failure?
  5. Why does an allowed code tool still need an isolation profile?
  6. Which memory fields should remain structured instead of being trusted from a compressed summary?

If one answer feels vague, return to that Lesson and rerun its Lab with the suggested change. The second half of the level assumes these boundaries are usable, because parallelism and delegation make mistakes harder to diagnose when the underlying state and recovery rules are unclear.

Check your reasoning with an applied failure

Suppose a worker sends a payment request, loses its connection, and crashes before recording completion.

A safe recovery path keeps the task ID, workflow state, operation ID, and prior attempt evidence in durable storage. After restart, the harness should reconcile the operation before issuing another payment. A queue may redeliver the task, but redelivery does not make a duplicate side effect acceptable.

Use the same case to test the six questions above:

  • durable state must contain the facts needed to resume safely;
  • a stable operation ID lets the system ask whether the earlier effect already happened;
  • trace fields should separate a policy denial from a tool/runtime failure;
  • an allowed code tool still needs isolation because permission to call it is not permission to reach everything on the host;
  • security policy, sandbox limits, and memory policy solve different problems;
  • important identity and authority fields should stay structured rather than being trusted only from a compressed summary.

Next Lesson

Continue with L12.7 — Parallel Tool Calls.

Lesson actions

Completion is stored locally on this device.

View progress