An agent run can contain several model calls, branching tool requests, retries, and work that continues after the original interaction ends. A transcript alone cannot reliably explain that execution. To investigate it, you need a record of what the system attempted, which checks allowed the action, what the tool reported, and how the application established its final outcome.
Agent logging works best when it follows the structure of the work. Give the whole run an identity, give each logical step an identity, and distinguish individual attempts. This guide develops that structure with practical examples. The agent logging hub summarizes the main event boundaries and recommended fields.
Use a trace to represent related operations
The OpenTelemetry explanation of traces describes spans as units of work with timing, context, and attributes. Related spans can share a trace identifier and use parent relationships to describe sub-operations. Apply that idea to an agent run: the run contains steps, while a step may contain model invocations, validation, and tool execution.
Trace structure supplies execution relationships; application events supply their meaning. A tool span lasting two seconds does not by itself explain whether the tool returned valid data, performed a requested change, or produced a result the agent ignored. Add bounded outcome fields and events for the decisions that matter.
Choose stable operation names such as “document.lookup” or “record.update.” Keep a customer’s name, document content, or full command out of the span name. Stable names also make aggregate analysis easier.
Separate runs, steps, calls, and attempts
A run is the overall application task. A step is a logical unit inside that task. A tool call is the system’s request to perform a particular operation. An attempt is one execution of that request. Document these definitions because frameworks may use the same words differently.
When a call is retried, preserve the logical call identifier and create a new attempt identifier. When the agent revises its plan and makes a materially different request, create a new call identifier. That distinction helps an investigator tell recovery from new work. It also makes repeated actions visible without relying on timestamps or similar-looking argument text.
Keep an event identifier separate from all of these. Re-exporting the same event should retain its event identity, while performing the operation again should create a new attempt. Otherwise, a delivery duplicate can look like an extra tool execution.
Log the boundary around every consequential tool
A useful tool sequence records request creation, argument validation, authorization, execution, result validation, and the final application decision. Not every stage needs a separate record; combine stages when their relationship remains clear. The requirement is that an investigator can establish what happened at each meaningful boundary.
This fictional event illustrates a completed read operation. Its fields are suggested application conventions, not an operational endpoint or a framework standard.
{
"event_name": "agent.tool_attempt.completed",
"run_id": "run_example_12",
"step_id": "step_example_03",
"tool_call_id": "call_example_08",
"attempt_number": 1,
"tool_name": "document.lookup",
"authorization_outcome": "allowed",
"execution_outcome": "completed",
"result_validation": "accepted"
}
Log a tool’s stable name and version when the version affects behavior. Prefer a classified input description and an opaque resource reference over raw arguments. A lookup term can itself contain customer information, and an output object can contain more information than the agent needs to preserve.
Keep authorization evidence connected to execution
For an action with access or change implications, record which policy decision applied and which permission scope was evaluated. Use a safe policy identifier and decision reason. The fact that a model requested a tool does not establish that the application authorized its execution.
If a human review is part of the workflow, distinguish approval requested, approval received, approval declined, and approval expired. Bind a decision to the specific action and the version of its arguments that was reviewed. If the action changes afterward, the earlier review should remain associated with its original request.
Keep identity records appropriately protected. An opaque actor reference and actor type can be enough for routine operational analysis. The authentication and MFA audit-log guide explains related distinctions between an identity event, an access decision, and an application outcome.
Represent retries and uncertain outcomes honestly
A timeout tells you that one observer stopped waiting. It does not necessarily tell you whether a remote change occurred. Record the timeout and the last confirmed execution stage. For a consequential action, use an explicit unresolved outcome until the application can reconcile the remote state.
Where the destination supports an idempotency mechanism, use its documented behavior to reduce the risk of repeating a change. Record a safe reference to that mechanism and the retry relationship. Do not assume that repeating the same arguments is automatically safe, or that a logging identifier itself enforces idempotency.
Separate recovery from reconciliation
A recovery attempt tries to complete the task. Reconciliation determines what a previous uncertain attempt actually did. Record them as distinct operations. An illustrative update might time out, then a read operation checks whether the requested version exists. That check provides new evidence; it should not erase the fact that the original observer saw a timeout.
Bound retries with an application policy, and record the reason when the policy stops further attempts. A run ending because its attempt budget was exhausted needs a different outcome from a run stopped by an access decision.
Preserve context across queues and parallel work
When a step moves to a queue, carry the correlation context in the approved message metadata and record the queue boundary. The worker should retain enough context to connect its result to the originating run. Keep message identifiers and delivery attempts distinct from tool-call identifiers and execution attempts.
For parallel branches, preserve the parent relationship and a branch identifier. Record when a branch finishes and how the joining step uses its result. If the run ends while a branch remains active, indicate whether that branch was cancelled, detached deliberately, or left unresolved.
Span links can express relationships that do not fit a simple parent-child tree, including work resumed in a separate trace. Use the tracing system’s actual context APIs and verify the emitted relationships. Copying a trace identifier into an arbitrary field does not by itself create a valid distributed trace.
Measure time and usage at the right scope
Measure total run duration separately from individual step durations. Parallel steps overlap, so their durations cannot simply be added to reconstruct elapsed wall time. Queue waiting, approval waiting, model execution, and tool execution also describe different causes of delay and should remain distinguishable when they matter to the user experience.
Attach model usage to the attempt that consumed it, then define how run-level totals include those attempts. The LLM usage logging guide explains how to avoid duplicated counts and preserve accounting gaps. A run can have a known application outcome while some usage remains incomplete.
Record observable decisions without unnecessary content
Capture the selected tool, the validated action, the policy result, and a bounded reason code for application decisions. Preserve relevant configuration versions and safe evidence references. These records describe behavior that the application can verify without requiring unrestricted collection of model text.
Apply the same content controls to tool arguments, tool responses, exception details, and intermediate model outputs. The prompt privacy guide describes a metadata-first approach. Treat logging controls as part of the tool boundary so that one verbose integration does not silently bypass the rest of the system’s collection policy.
Define the final run record
End a run with an application outcome that explains whether the requested task was completed, partially completed, declined, cancelled, or left unresolved. Record the last confirmed step and any outstanding operation references. A tool’s successful response should influence this result only after the application has validated what that response means. If the user-facing answer reports completion while a consequential operation remains uncertain, the inconsistency should be visible in the record. Keep subsequent reconciliation attached to the run so reviewers can distinguish the original report from later evidence.
Conclusion: make a run reconstructable
A useful agent trace connects each run to its steps, attempts, permissions, and verified outcomes. Before relying on it, inspect synthetic cases covering retries, declined actions, parallel branches, and uncertain remote results. A reader should be able to identify the last confirmed state and the evidence behind the final decision. That is the practical standard for agent logs that support debugging and operational review.



