LLM Token Logging: Measure Usage, Latency, and Cost
Build an interpretable record of LLM usage, timing, retries, and cost estimates while preserving incomplete accounting states.
Track LLM token usage with clear provenance, accounting states, attempt identifiers, category definitions, and rate references.
A token measurement needs both a value and a meaning. Preflight estimates, provider-reported usage, and reconciled billing records answer different questions. Preserve the source of the observation and whether its coverage is complete. Missing usage should remain unknown rather than becoming zero so a chart can display a convenient total.
Document each provider-specific category before combining values. A category may already be included in a broader total, so blindly summing every available field can count the same work twice. Keep model identity and the applicable mapping version available when investigating a change in accounting.
A user request may create multiple model executions through retries, fallbacks, or agent steps. Record usage for each attempt and define how request-level totals include those attempts. Give each exported observation its own event identity so that delivery duplicates can be recognized separately from repeated model work.
Choose an explicit correction rule. A later observation may replace an earlier partial total, or it may represent a separate adjustment. Both approaches need stable identifiers and documented aggregation behavior. For interrupted streams, preserve the last observed state and the possibility that final usage remains unavailable.
Associate estimates with a versioned rate reference that specifies units, currency, model, and relevant usage categories. Preserve the reference used at calculation time so a later change does not silently rewrite historical interpretation. Reconcile estimates against the authoritative billing record before treating them as confirmed charges.
The token usage, latency, and cost guide covers accounting completeness, latency boundaries, and a small duplicate-delivery exercise. The recommended fields below are a starting point for an interpretable usage contract, not a claim that every provider reports identical data.
Illustrative field suggestions for your own event contract. Adapt the names, values, and collection rules to your system.
| Field | What it helps explain |
|---|---|
attempt_id | Tie usage to the specific model execution. |
usage_source | Distinguish estimates from observed or reconciled records. |
usage_state | Make complete, partial, and missing coverage explicit. |
input_tokens | Store input usage under a documented provider mapping. |
output_tokens | Store output usage without mixing accounting categories. |
rate_reference | Reference the dated assumptions behind a cost estimate. |