AI & LLM / THE SIGNAL

LLM Token Logging: Measure Usage, Latency, and Cost

Build an interpretable record of LLM usage, timing, retries, and cost estimates while preserving incomplete accounting states.

Count the Tokens neon card with translucent cyan and lime token stacks.

LLM usage becomes difficult to explain when one “request” can include a long conversation, retrieved documents, several model attempts, and multiple tool calls. A single token total does not tell you which work produced it. A single duration does not explain where a user waited. Useful token logging connects usage and timing to the operation that consumed the resources, then preserves enough context to compare like with like.

The goal is an accounting record that supports engineering decisions. You should be able to investigate growing prompts, repeated retries, incomplete streams, and unexpected allocation changes without retaining the actual conversation. Begin with the terminology in the LLM logging hub, then define the measurements your application can obtain reliably.

Keep estimates, observed usage, and charges separate

A preflight token estimate answers whether a proposed input appears to fit a chosen budget. Observed usage describes what a provider reports for a completed or partially completed operation. An invoiced charge describes the commercial accounting of that work. These are related measurements with different purposes and different uncertainty.

For example, Anthropic’s token-counting documentation describes token counting before message creation and explicitly characterizes its result as an estimate that can differ from actual input usage. This supports a practical rule: label estimates as estimates, and preserve observed values separately when they become available. Do not silently overwrite one with the other.

Use a provenance field such as usage_source and an availability field such as usage_state. Their allowed values should come from your documented contract. Missing usage is not zero usage. A timeout can leave accounting unresolved even when the application has already reported a failure.

Choose the right accounting unit

Record usage against an individual model attempt, and connect attempts to an overall request or run. If a summarization job retries once, retain both attempts. The final answer may come from the second attempt, while both attempts may matter when reconciling usage. A fallback to a different model also needs a distinct record and the identity of the model actually used.

Decide whether a usage record is an immutable final observation or an update that replaces an earlier observation. Both designs can work. Mixing them produces double counting. If updates replace earlier values, include an observation version and a stable accounting key. If records are immutable, represent corrections explicitly and define how aggregation applies them.

The token accounting topic hub collects the recommended fields and boundary questions that help keep this contract understandable.

Use a narrow, typed usage event

The following record illustrates a possible application-level shape. Its identifiers and values are fictional, and its names are not a universal provider schema.

{
  "event_name": "model.usage.observed",
  "request_id": "req_example_04",
  "attempt_id": "attempt_example_02",
  "usage_source": "provider_response",
  "usage_state": "complete",
  "input_tokens": 1200,
  "output_tokens": 180,
  "duration_ms": 2400,
  "accounting_version": 1
}

Store counts as integers and timing in a declared unit. Preserve provider-specific categories when they affect interpretation, but document whether each category is included in another total. A cached-input field, for example, must not be added to a total that already contains it. Only calculate a normalized total after validating the selected provider’s definitions.

Make the model identity precise enough to compare

Keep a safe model identifier, provider identifier, and application configuration version. If a routing alias can resolve to different models, record the requested alias and observed model separately when available. An aggregate labeled only “default” becomes hard to interpret after its routing configuration changes.

Handle streams and interruptions explicitly

For a streaming response, establish when usage becomes authoritative for your integration. Some information may arrive only near completion. Buffer the accounting state until the integration knows which observation is final. Do not assume every chunk carries an independent count, and do not sum cumulative counters as though they were deltas.

Track the stream outcome independently: completed, interrupted, cancelled, or unknown are useful starting concepts. If a stream ends before the application receives final accounting, preserve the partial evidence and mark its coverage. A downstream connection failure should not turn an unavailable usage field into zero merely to keep a dashboard calculation simple.

Test this behavior with synthetic interrupted streams. The useful question is whether a reader can tell what is known, what is incomplete, and whether a later observation should replace or supplement the record.

Measure latency where users experience it

Define the boundaries of each duration. Application latency might begin when a task is accepted and end when its result is usable. Model-call latency might begin immediately before an invocation and end after the response is processed. Queue time, retrieval, retries, and output validation can explain the difference between those measurements.

For a streaming experience, record the time to the first meaningful output separately from total completion time, if your integration exposes those moments. State precisely whether that first output means any received event, the first text fragment, or the first material displayed to a user. Otherwise, implementations can use the same metric name for different experiences.

Use a monotonic clock for elapsed duration inside one process when available. Wall-clock timestamps help correlate records, but clock adjustments can distort a duration calculated by subtracting them. For work crossing process boundaries, preserve the measurement location and avoid pretending that independently measured clocks form a perfectly synchronized stopwatch.

Calculate cost from a versioned rate reference

Keep the usage event independent of a current rate table. A cost estimate can reference a separate, dated rate definition with an explicit currency, unit, model, and relevant usage category. This allows an investigator to explain how an estimate was calculated without placing commercial terms or mutable configuration into every event.

For each non-overlapping category, divide the token count by the rate’s token unit, multiply by the applicable rate, and then sum the results. Account separately for any relevant charge that is not token-based. Preserve the assumptions and rate-reference version. Label the result as an estimate until it has been reconciled with the authoritative billing record.

Do not revise historical estimates merely because today’s rate reference changes. If a correction is necessary, keep its reason and effective period. This makes comparisons reproducible and prevents an unexplained shift in a historical dashboard.

Compare workloads before comparing totals

Group analysis by a small set of meaningful dimensions: workflow, model, prompt version, outcome, and deployment environment. Separate experimental runs from production traffic. A batch of long-document summaries should not be evaluated against short chat replies without accounting for the different work each performs.

Inspect distributions as well as totals. A rising average may reflect a few unusually long inputs or a broad increase across all requests. Examine request counts, typical input sizes, upper-range durations, and retry frequency together. Token totals alone cannot tell whether the application grew, its workload changed, or a defect caused repeated work.

When prompt changes are involved, connect the measurement to a version rather than the prompt text. The prompt versioning guide explains how to preserve that reference while keeping user content out of routine logs.

Make accounting completeness visible

A usage dashboard should state what fraction of operations has complete accounting. Count missing observations, unresolved attempts, and rejected usage records alongside recorded totals. If a deployment changes the usage parser, compare coverage before interpreting a sudden reduction as an efficiency improvement.

Test one clean success, one retry, one stream interruption, and one duplicated export. Verify that a request-level total includes the intended attempts exactly once. For multi-step workflows, follow the agent tracing guide to connect model usage to the tool sequence that caused it.

Audit a small ledger before trusting a large total

Consider a fictional request with two model attempts and three exported usage records because the second record was delivered twice. A correct aggregation should recognize two attempts and one delivery duplicate. If the first attempt has incomplete accounting, the request total should also expose that gap. Write the expected result down before inspecting the dashboard. Then introduce a corrected observation and confirm that it follows the contract’s replacement or adjustment rule. This small exercise can reveal double counting and silent missing values more clearly than a large aggregate.

Conclusion: preserve meaning alongside the numbers

Reliable LLM token logging gives every count a source, every duration a boundary, and every cost estimate a reproducible set of assumptions. Keep attempt-level records, preserve incomplete states, and validate aggregation with realistic failure paths. Those choices make usage trends explainable without collecting the sensitive conversations behind them.

You’ve reached the end of this field note.
Keep exploring The Signal