본문으로 건너뛰기
L12.2

Task Queues and Durable State

Goal

Model agent work as durable tasks that can wait in a queue, be claimed by workers, and resume from recorded state after interruption.

Separate the task from the worker​

Imagine a teacher has a box of assignment cards and several helpers. A card can wait in the box until a helper is free. If one helper goes home halfway through a job, the card and the written progress still exist, so another helper can continue. The work does not disappear with the person who was doing it.

A task queue and durable state solve the same separation in software. The queue says which work is waiting to be handled. Durable state records facts that must survive a worker or process disappearing. A worker is temporary; the task record is the recoverable source of truth. This is why "the Python variable still exists" is not enough evidence that a long-running job can recover.

A queue is a controlled place where work waits until a worker can handle it. The queue does not need to contain the entire conversation or every tool result. It usually needs enough identity to find the durable task record. That record stores the task's current state, relevant inputs, completed operations, retry metadata, and the next legal transition.

Claim and recover work safely​

Imagine task T-42 is waiting to check an order. Worker A claims it, records RUNNING, calls a read-only tool, and then crashes before producing the final answer. Worker B can later claim T-42. If the read result or completed step was recorded durably, Worker B can continue. If nothing was recorded, it must repeat work or guess.

T-42: QUEUED → RUNNING → [worker crash] → RUNNING → DONE
durable task identity stays the same

Claiming work introduces concurrency questions. Two workers must not both believe they exclusively own the same task. Systems solve this with leases, locks, compare-and-set updates, or workflow-engine semantics. You do not need one particular technology to understand the rule: task ownership must be represented by application data or a service that provides equivalent guarantees.

Durable state should be compact and purposeful. Storing every transient Python object makes recovery brittle. Prefer stable data such as task ID, state name, operation IDs, inputs needed for replay, outputs that matter later, timestamps, and version identifiers. Temporary caches can be rebuilt; authority and completion facts should not be guessed.

Make transitions and backpressure explicit​

A state transition should be explicit. For example, QUEUED may move to RUNNING, RUNNING to WAITING_APPROVAL, and WAITING_APPROVAL to QUEUED after approval arrives. Writing these transitions down prevents a restarted worker from inventing a shortcut directly from WAITING_APPROVAL to COMPLETED.

Queues also help with backpressure. Backpressure means slowing the rate at which work begins when downstream capacity is limited. Without it, a burst of requests can start thousands of expensive agent runs at once. A queue lets the system control concurrency and observe how much work is waiting.

A queue does not make retries safe​

One misconception is that a queue itself makes operations safe to repeat. It does not. Queues can redeliver messages, workers can crash, and network acknowledgements can be lost. The next lesson introduces retries and idempotency because durable task delivery and duplicate-safe side effects are separate problems.

For this level, focus on the data model rather than a specific hosted queue. A deterministic list of task records is enough to practice claim, transition, crash, and resume behavior. The same reasoning carries into workflow systems such as Temporal, whose documentation describes durable execution and workflow state.

Predict

A worker crashes after marking a task RUNNING. What should allow another worker to continue safely?

Run the local Lab​

Run:

python labs/notebooks/level-12/l12-02-task-queue.py

The script puts one task in a queue. Worker A claims it with a lease, then crashes. Worker B keeps trying to claim the same task.

  1. Run it unchanged. With lease_steps = 2, read step 0: A claimed ... 'lease_until': 2.
  2. Read step 1: B blocked (A's lease runs until step 2) and then step 2: B recovered ... 'owner': 'B'.
  3. Change lease_steps = 2 to lease_steps = 1 and rerun.
  4. Now A's lease ends at step 1, and the output shows step 1: B recovered. In both runs, task_id stays 'T-42'.

A shorter lease recovers a crashed task sooner. But if the lease is too short, a slow but healthy worker could lose its task while it is still working. The task's identity never changes, only its owner.

Loading lab…

Quick Check

1. Why separate task lifetime from worker lifetime?
2. What does a queue most directly provide?
3. Which state is best kept durably?

0 of 3 questions answered.

Explain it back​

Explain how QUEUED, RUNNING, WAITING_APPROVAL, and COMPLETED could form a legal state machine. Include what should happen if a worker holding RUNNING disappears.

Key Takeaways

  • Durable tasks outlive individual workers.
  • Queues coordinate pending work and support backpressure.
  • Task claims need explicit ownership or lease semantics.
  • Durable records should contain recovery-critical facts, not every transient object.
  • Queues do not by themselves make side effects duplicate-safe.

Next Lesson

Next, L12.3 — Retries, Idempotency, and Recovery addresses what happens when a task or tool operation must be attempted again.

References

Lesson actions

Completion is stored locally on this device.

View progress