Tracing Agent Decisions
Goal
Create structured traces that connect a task, controller decisions, tool calls, retries, and terminal outcomes without recording hidden reasoning.
Trace the journey, not hidden reasoning
A package-tracking page is useful because it does not only say "delivery failed." It shows a sequence: package accepted, sorting center reached, truck departed, delivery attempted, and so on. Each event has a time and identity, so you can locate the stage where the problem appeared.
A software trace serves a similar purpose for one agent run. The trace links the whole journey, while smaller spans record timed pieces of work inside it. Different identifiers answer different questions: a task ID names the durable job, a trace ID links one observed run, a span ID names one operation inside that run, and an operation ID can connect retries of the same side effect. You do not need to memorize the names first; ask what relationship each ID lets you reconstruct.
A trace is a set of related records for one request or task. A span represents one timed unit of work inside that trace. For an agent harness, useful spans might include task execution, model request, policy check, tool call, memory retrieval, approval wait, and recovery step. Parent-child relationships show which operation caused another operation.
Use structured fields and distinct IDs
Structured fields make traces searchable. Instead of a log line that says 'tool failed again,' record fields such as task_id, action, tool_name, attempt, error_type, duration_ms, and terminal_reason. Consistent fields let you compare many runs and answer questions such as which tool creates the most retries.
{"task_id":"T-42","trace_id":"tr-9","span":"tool_call","attempt":2,"duration_ms":140}
Tracing should record application-visible decisions, not hidden chain-of-thought. The harness can record that the model proposed search_catalog, that policy allowed it, that the tool returned zero items, and that the controller chose a second search. That is enough to debug behavior without storing private internal reasoning text.
Identifiers matter. A trace ID links the whole run. A span ID names one operation. A task ID links events across restarts or multiple traces. An operation ID links retries of the same side effect. These identifiers answer different questions, so collapsing them into one field makes recovery and analysis harder.
Timing and privacy are part of the trace
Timing is also evidence. End-to-end latency can be broken into queue wait, model latency, tool latency, approval wait, retry delay, and local processing. If the total run takes thirty seconds, tracing lets you see whether the time was spent waiting for a user, calling a slow service, or repeatedly retrying one tool.
Be deliberate about what you record. Tool arguments may contain secrets or personal data. Traces should apply redaction and retention rules. Observability is not an excuse to copy every prompt, credential, or user field into a permanent log.
OpenTelemetry provides common concepts and data models for traces, metrics, and logs. The exact SDK may change, but the useful mental model is stable: emit structured telemetry at meaningful boundaries and preserve relationships between operations.
Record evidence that answers questions
A trace is valuable only if it helps answer operational questions. Before adding a field, ask what decision it enables. Before omitting a field, ask whether you could still distinguish a model-selection problem from a policy rejection, tool failure, or retry bug.