Skip to main content
L12.11

Cost, Latency, and Reliability

Goal

Measure cost, latency, and reliability together and identify the critical path so optimizations do not hide safety or quality regressions.

Track more than one objective​

Suppose two delivery plans both succeed. Plan A uses one truck and arrives in 20 minutes. Plan B uses four separate trips and arrives in 18 minutes. Calling B simply "better" hides extra cost and more opportunities for something to go wrong.

Agent operations have the same multi-objective tradeoff. A design can reduce latency by running more work in parallel while increasing tool cost, resource pressure, or failure surface. Another design can retry aggressively and raise success on transient failures while making slow cases much slower. Measure cost, latency, and reliability together so an optimization does not quietly move the problem into another column.

Break latency into causes​

Latency is the time a user or downstream system waits. Break it into components: queue wait, model time, tool time, retry delay, approval wait, and local processing. This decomposition matters because the best fix depends on where time is actually spent.

Parallel work changes how latency is calculated. If two independent calls take 200 ms and 300 ms and run together, their contribution to the critical path is closer to 300 ms than 500 ms. Cost may still include both calls. This is why latency and cost should not be treated as the same quantity.

sequential: 200 ms + 300 ms = 500 ms
parallel independent calls: max(200, 300) = 300 ms

Reliability includes more than whether the final answer exists. Useful metrics include task completion, recovery success, duplicate-effect rate, timeout rate, policy violations, and fraction of runs stopped by budget. The harness should expose which failure category changed after an optimization.

Compare tradeoffs under budgets​

Suppose a cheaper model reduces cost by 40% but increases retries enough that end-to-end latency rises and task success falls. The cheaper request price did not produce a better system. Operational decisions should compare a vector of metrics rather than a single headline number.

Budgets can be enforced per task. A controller can track maximum model calls, tool calls, elapsed time, and estimated cost. Reaching a budget should create an explicit terminal reason rather than silently producing lower-quality behavior.

Service-level objectives turn measurements into targets. An objective might say that 95% of read-only support tasks should finish within five seconds while unauthorized execution remains zero. The exact target is application-specific; the important point is that performance and safety expectations are explicit.

Use percentiles and traces​

Percentiles help describe latency distributions. An average can look good while a small group of users waits much longer. p95 latency asks for a value that 95% of measured runs finish at or below. You do not need advanced statistics here, but you should avoid reporting only the mean.

Use traces to connect metrics back to causes. If p95 latency worsens, inspect whether queue wait, tool latency, or retries changed. A metric alerts you to a problem; traces help explain where it came from.

Predict

Two independent calls take 200 ms and 300 ms and run fully in parallel. Ignoring overhead, what is their approximate contribution to critical-path latency?

Run the Docker-environment Lab preflight​

The full activity is registered for the Docker-oriented environment used by later operational work. Start with this deterministic Python preflight from your repository checkout so you can separate controller-logic failures from Docker, cloud, GPU, or credential setup problems.

Run:

python labs/notebooks/level-12/l12-11-cost-latency.py

The Lab calculates total cost, sequential latency, parallel critical-path latency, and p95 over fixed traces. Change one tool duration from 120 to 900 and rerun. Observe how p95 and the critical path change while unrelated call costs stay the same.

Loading lab…

Quick Check

1. Why measure cost, latency, and reliability together?
2. What does p95 latency describe?
3. What should happen when a task budget is exhausted?

0 of 3 questions answered.

Explain it back​

Given a run with queue wait, two parallel reads, one model call, and one retry, explain how you would calculate critical-path latency and total operation cost separately.

Key Takeaways

  • Operational quality has several dimensions at once.
  • Break latency into meaningful components.
  • Parallelism reduces critical-path time only for independent work.
  • Use percentiles to see slow-tail behavior.
  • Budget exhaustion should be explicit and measurable.

Next Lesson

Next, L12.12 — Operating an Agent in Production combines queueing, recovery, tracing, isolation, and release controls into an operational routine.

References

Lesson actions

Completion is stored locally on this device.

View progress