A token logger for a language-model application records usage quantities and the context needed to interpret them. Its useful output is a compact record of which operation ran, which model was requested, what usage was reported, and where information is missing. That record can support investigation and planning without storing the messages being processed.

The first design decision is what you intend to count. A user action, a model call, a network attempt, and a delivered answer can all have different totals. This guide proposes an application accounting record with explicit boundaries. The Token Logger overview introduces the terms and the role of usage records within a logging system.

Start with the accounting question

Choose a question such as “how much reported usage belongs to this feature each day?” Then define the feature, the time boundary, and the records included. A drafting feature might include one model call in the normal path and additional calls when validation fails. Counting only delivered answers would hide that difference.

Keep a short accounting note alongside the event schema. State whether a row represents one observed attempt, one logical model operation, or a provider usage report. Define the reporting timezone and the rule for operations crossing a daily boundary. These are product decisions, but inconsistent answers can make two correct queries produce different totals.

Design a record with provenance

Use a small, explicit field list. Provenance means preserving where a value came from, so a later reader can distinguish reported quantities from estimates and corrections. The following fields describe an illustrative application record rather than a prescribed telemetry standard.

FieldPurpose
operation_idIdentify the logical work being counted.
usage_record_idRecognize repeated delivery of the same usage record.
featureGroup work using a controlled application label.
requested_modelPreserve the configured model or route.
usage_sourceDistinguish provider reporting, local estimation, and reconciliation.
usage_statusExplain whether usage is complete, partial, or unavailable.

Add input and output quantities when they are available, and retain the meaning supplied by the provider. If the provider distinguishes consumed and billable quantities, document which one the application record contains. Store the relevant model identity when it is actually returned; an unresolved route should remain unresolved.

Make unknown values visible

A missing count is not evidence that an operation used zero tokens. In the proposed application record, retain a missing value together with an explanation. For example, a client may have observed an interrupted operation without receiving final usage information. That case deserves a visible accounting gap.

{
  "usage_record_id": "usage-example",
  "operation_id": "operation-example",
  "feature": "document-summary",
  "usage_source": "provider_response",
  "usage_status": "unavailable",
  "input_tokens": null,
  "output_tokens": null,
  "reason": "final_usage_not_observed"
}

These identifiers are fictional. The example makes no claim about any provider’s billing behavior. If you later recover an authoritative usage value, write a correction that refers to the original record, or use a documented replacement rule. Silent overwrites make a past report difficult to reproduce.

Understand totals and detailed categories

The OpenTelemetry inference token metrics conventions distinguish consumption counters from per-operation distributions and describe cache and reasoning categories as subsets of broader totals. The document is marked Development. Review and pin the conventions used by your exporter; an application accounting record can preserve additional completeness information separately from the metric representation.

Draw the relationship between totals and categories before writing an aggregation. Suppose an invented provider report contains 1,000 input tokens, of which 300 are identified as cached. If the cached count is included in the input total, adding both produces an incorrect 1,300-token total. Preserve the total and the detail as different views of the same operation.

Apply the same discipline to modalities. If text, image, or audio quantities are available, document whether they partition a total or overlap with other categories. When that relationship is undocumented, preserve the supplied fields and mark the unresolved interpretation. A tempting sum is not a substitute for a definition.

Keep retries from distorting the result

Separate the identity of a logical operation from the identity of each observed attempt. One operation may have multiple attempts; one usage record may also be delivered more than once. Those situations require different treatment. A retry can represent additional work, while duplicated delivery can represent the same evidence arriving twice.

Choose a stable deduplication key at the point where usage becomes an application record. Keep the record’s identity unchanged when a collector retries delivery. Test whether your storage or aggregation layer recognizes that duplicate. Do not deduplicate solely by matching token counts: two legitimate requests can report identical quantities.

If a framework reports only the final logical operation, retain that boundary in the schema. Avoid inventing attempt-level usage from a total. The AI Logging guide explains how operation boundaries and retry observations fit into the wider workflow.

Estimate cost with a versioned calculation

Keep usage quantities separate from estimated money values. A calculation needs an identified price schedule, model mapping, currency, unit size, and effective period. Store the revision of those assumptions so the same usage can be recalculated when an interpretation changes without rewriting the original evidence.

For a deliberately simplified example, imagine an application with only input and output billing categories. Its estimate is input quantity multiplied by the configured input rate, plus output quantity multiplied by the configured output rate. Convert both quantities to the rate’s unit before multiplication. If rates are per million tokens, divide the corresponding counts by one million first.

Real arrangements may require additional categories or adjustments, so label the output as an estimate until reconciled with the applicable usage and billing records. Do not treat missing usage as a free operation. Show the count of unresolved records beside an estimate so a reader can judge its completeness. Broader infrastructure tradeoffs are covered in logging storage and cost planning.

Choose groupings that stay useful

Define a controlled set of feature labels such as document summary, classification, or answer drafting. A label should describe a durable application function, not contain user text. Stable labels make it easier to compare equivalent work across releases and to assign an owner when behavior changes.

Keep individual operation identifiers available for investigation without making every identifier a permanent metric grouping. Decide which questions need an aggregate and which need a record lookup. A daily feature total and an individual request investigation have different access patterns; they do not need identical representations.

Compare equivalent workloads

Before interpreting a rise in average usage, check whether the mix of work changed. A feature processing longer documents may use more input even when its implementation is unchanged. Compare the same feature and configuration revision, and separate call counts from quantities per call. Record a safe size category when that helps explain the workload without preserving document content.

Keep the question attached to the chart. A total answers how much reported usage accumulated; a per-operation distribution helps locate unusually large requests. Each view becomes more useful when the reader knows its scope and completeness.

Keep secrets and content outside usage events

Build the usage event from approved fields instead of serializing a response object wholesale. Review model names, route labels, exception messages, and custom metadata as well as the obvious content fields. A label created from a document title may reveal information even when the document itself is absent.

Use synthetic sentinel values in tests to represent messages, authorization values, and private identifiers. Exercise normal responses and failures, then inspect the exported records for those sentinels. Make the logging helper the common path for application code so its field policy remains reviewable. The redaction and retention guide covers access and deletion decisions around operational data.

Validate the accounting before trusting a chart

Create a small fixture set whose expected result you can calculate by hand. Include a complete record, missing usage, a duplicate record, a retried operation, and a total with a documented subset. Run the records through the same processing path used for real traffic, then compare the stored results with your expected table.

Check corrections and daily boundaries too. A late record should follow your chosen reporting rule instead of appearing unpredictably in whichever day a query happens to run. Keep the number of incomplete or rejected records visible. A chart that displays only accepted quantities can otherwise conceal a failing collection path.

A useful token logger preserves both quantities and their interpretation. Start with clear accounting boundaries, explicit provenance, and controlled metadata. Keep uncertainty visible, verify arithmetic with small examples, and use reconciliation to improve the record when better evidence becomes available.