LLM Token Logging: Measure Usage, Latency, and Cost
Build an interpretable record of LLM usage, timing, retries, and cost estimates while preserving incomplete accounting states.
Learn how to observe LLM calls, model routing, response validation, interrupted streams, and the difference between requests and attempts.
An LLM call sits inside a larger application task. The task may assemble a conversation, retrieve context, select a model, validate a response, and retry after a failure. Give each attempt its own identity while preserving the connection to the overall request. This makes a recovered error visible without confusing it with a failed user task.
Record the model identifier and configuration version available to the application. When a mutable routing alias selects the model, preserve its resolved identity when the integration exposes it. Historical comparisons become difficult if every record simply names the model “default.”
A response can arrive successfully and still fail the application’s rules. Record validation outcomes for the format and required content your feature expects. Use bounded reason codes for missing fields, unsupported structure, or other defined checks. Avoid placing the full rejected response inside an exception log merely to make the error easy to read.
Streaming calls also need a clear completion boundary. An interrupted stream may have delivered some text while leaving its final usage or status unknown. Preserve the observed state and the reason the application stopped waiting. Do not imply that cancellation from one observer establishes the outcome of all remote processing.
Compare calls from similar workflows and prompt versions. Separate queue delay, model-call duration, and application validation when those distinctions help explain the user experience. Attach usage to the corresponding attempt so that a later retry or correction does not silently change which work a number represents.
The LLM usage and latency guide develops those accounting boundaries in detail. The fields below are recommended design examples; add provider-specific metadata only after its meaning and collection purpose are documented.
Illustrative field suggestions for your own event contract. Adapt the names, values, and collection rules to your system.
| Field | What it helps explain |
|---|---|
request_id | Connect the attempt to its application request. |
attempt_id | Identify one model execution attempt. |
model_id | Record the safe model identity available to the application. |
prompt_version | Reference the resolved input configuration. |
response_status | Preserve the observed response or stream state. |
validation_code | Explain the application’s bounded output check result. |