A useful AI log explains what happened around a model interaction well enough to investigate it later. It connects the application request, the selected workflow, the model attempt, and the result the application accepted. It does not need to preserve every message or turn a troubleshooting system into a second copy of customer data. The design starts with decisions: what would an engineer need to know when a response is slow, incomplete, rejected, or unexpectedly expensive?
This guide develops a practical event contract for those questions. The field names and examples are illustrative application conventions to adapt to an actual implementation. The AI logging topic hub provides a starting point for related terminology and recommended fields.
Start with the questions your events must answer
Write down three investigations before writing instrumentation. For a slow response, you may need to separate queue time, retrieval time, model time, and application processing. For an invalid response, you need the prompt version, output contract, validation result, and fallback decision. For a sudden usage increase, you need the workflow, number of attempts, and usage reported for each attempt.
Turn those questions into an event inventory. An event should represent an observable change or completed operation: a request was accepted, retrieval finished, a model attempt ended, or an output failed validation. Avoid messages such as “AI working” that lack a stable meaning. Give each event an owner who can explain when it fires and what its fields mean.
Separate transport success from application success
A successful network response can still contain output your application cannot use. Record the transport result and application validation result separately. If an answer passes a JSON parser but omits a required business field, the event should preserve that distinction. Otherwise, dashboards can report success while users repeatedly encounter failures.
Build a small, consistent event envelope
Use a common envelope across workflows before adding model-specific attributes. The OpenTelemetry Logs Data Model distinguishes event time from observation time and provides places for severity, trace context, resource information, and attributes. That distinction is useful when records arrive late or move through several collection stages. Align an implementation with the relevant specification instead of assuming a handwritten JSON object is already compliant.
For an application-level design, begin with an event name, schema version, event identifier, timestamp, service, environment, and request identifier. Add an outcome with a documented vocabulary. Keep values typed consistently: durations as numbers in a named unit, counts as integers, and missing values as explicitly unknown rather than arbitrary text.
{
"event_name": "model.attempt.completed",
"schema_version": 1,
"event_id": "evt_example_01",
"request_id": "req_example_01",
"attempt_number": 1,
"workflow": "document_summary",
"outcome": "validated",
"prompt_version": "summary-v3"
}
This simplified example describes a completed attempt without including a document, prompt, response, or secret. A real contract also needs time, source, and collection fields appropriate to the chosen telemetry system.
Describe the full request lifecycle
A practical lifecycle includes acceptance, preparation, execution, validation, and completion. Record each boundary only when it helps reconstruct behavior. Preparation might involve selecting a prompt template and retrieving documents. Execution might involve one model call or several. Validation decides whether the returned material satisfies the application’s rules. Completion describes what the application ultimately delivered.
Give the overall request its own outcome. A first attempt may fail and a fallback may succeed; both facts matter. If the only record says “success,” engineers lose the retry history. If the only record says “error,” they may mistake a recovered attempt for a failed user task. Use an overall request identifier to join the lifecycle and distinct attempt identifiers for individual executions.
Include cancellation and timeout paths in the design. A client disconnect does not necessarily prove that remote processing stopped. Where the remote outcome cannot be established, preserve an unknown state and allow a later reconciliation event to clarify it.
Normalize carefully across providers and workflows
A shared schema makes comparison easier only when shared fields retain the same meaning. Define what “duration” measures, when an attempt begins, and whether usage values are estimates or observed totals. Keep original provider field names in a documented mapping outside ordinary event bodies, and retain only the bounded metadata needed to investigate mapping errors.
Do not force every provider-specific field into a generic total. Unsupported values should remain unavailable. Different request types may report different usage categories, and a field called “tokens” can hide several accounting choices. The guide to LLM token logging explains how to preserve those distinctions without confusing estimates, usage, and cost.
Choose identifiers for correlation, not exposure
Use opaque identifiers to connect related events. A request identifier should not contain an email address, document title, phone number, or prompt fragment. Keep the identifier’s scope clear: one user interaction, one background job, or one model attempt. Reusing a single session identifier for every operation makes individual investigations harder and can expose an unnecessarily broad history.
Separate high-cardinality identifiers from labels used for aggregate metrics. A request ID is valuable for finding one event sequence, but grouping every measurement by request ID creates a different series for each request. Use bounded dimensions such as workflow and outcome for summaries, then follow identifiers into detailed records when an investigation requires it.
Make content capture a deliberate exception
Start by recording metadata that answers the intended questions. Prompt template versions, input size, output format, validation codes, and redaction outcomes can often explain a failure without storing customer text. If a particular investigation requires content, define its purpose, collection scope, access rules, and deletion plan before enabling capture.
Review hidden copies as carefully as the main event body. Exceptions, debug middleware, request headers, retriever results, and serialized tool arguments can introduce data that an application developer never intentionally logged. The prompt logging and redaction guide describes a version-based approach that keeps ordinary operational records useful while reducing unnecessary content exposure.
Design for delayed, duplicated, and missing events
Treat collection as its own system with observable failure modes. A process can stop before a buffer flushes. A collector can reject an oversized record. A retry can deliver the same event twice. Decide which events may be dropped under pressure, which need durable buffering, and how the application should behave when logging cannot proceed.
Use an event identifier to recognize retransmission of the same record. A retried model operation is a new attempt, while a retried export of an existing event is a delivery duplicate. Preserve that difference. Reconstruct sequences with explicit relationships and timestamps rather than assuming records arrive in execution order.
Monitor the collection path with bounded counters for rejected records, export failures, queue age, and dropped batches. Keep those signals independent enough to reveal a broken log pipeline. A silent collector can otherwise make a service look healthy simply because its error records stopped arriving.
Review events with an investigation exercise
Before expanding instrumentation, walk through a small set of synthetic cases: ordinary success, invalid output, a recovered retry, cancellation, and collection failure. Ask someone unfamiliar with the implementation to reconstruct each outcome using only the emitted metadata. Any answer that requires undocumented assumptions identifies a gap in the event contract.
Then inspect the data itself. Check that field types match the contract, identifiers connect the intended operations, sensitive values are absent, and timestamps use the declared format. Review a proposed schema change against existing readers. Adding a field is usually easier to absorb than silently changing the unit or meaning of a field already used in alerts.
Follow one concrete failure through the record
Imagine a summary request that reaches the model successfully but fails the application’s required-field check. Start with the request outcome, follow its attempt identifier, and inspect the validation code and prompt version. Compare the same workflow before and after that version changed. If transport results remain stable while one validation code rises, investigate the prompt and output contract before changing the collection infrastructure. This sequence is an illustrative investigation, not proof of a particular cause. The record should help narrow the next question and preserve enough context to test it.
Conclusion: collect evidence that supports a decision
Strong AI logs connect a request to its attempts, decisions, and final outcome through a small, documented event contract. Begin with the investigations that matter, preserve uncertainty, and make collection failures visible. Expand only when a new field answers a real question. This produces an operational record that engineers can interpret consistently as workflows, providers, and application requirements evolve.



