AI logging becomes easier to design when you follow one model request through the application. A user action might trigger retrieval, a model call, a tool execution, another model call, and a final formatting step. If the only record says “AI request completed,” a slow response or failed workflow can remain difficult to explain.

A useful starting point is to record the boundaries, outcomes, and configuration of those operations. Add content capture only for a clearly defined need. This guide proposes a practical event design for developers who want to understand model-powered workflows. The AI Logging topic page places this approach alongside the other operational signals you may already collect.

Draw the workflow you actually own

Start with the application boundary: what action begins the work, and what observable event means the user-facing task has ended? Then list the operations in between. A support assistant might accept a question, retrieve reference material, request a draft, validate its format, and deliver the answer. Give each step a clear name that can stay stable across deployments.

Record what your application can observe. A client can measure when it sends a request and receives a response. That does not reveal every internal stage at the model provider. Keep this distinction in field descriptions so a client duration is not accidentally presented as server computation time. Draw unknown stages as unknown; filling them with inferred timestamps creates misleading precision.

Choose a question for each signal

Investigation questionUseful recorded context
Which step failed?Workflow identifier, operation name, outcome, and error category.
Where did the user wait?Explicit timing boundaries for retrieval, model calls, and delivery.
What changed between deployments?Application release and configuration revision.
Why did usage rise?Observed model calls, attempts, and available usage metadata.
Did the answer meet the task?A separate, defined evaluation result and its method.

Keep these questions separate in analysis. A completed network request and a useful answer are different observations. An operational event can establish that a response arrived, while a task evaluation checks whether the response met specified criteria. Naming those outcomes precisely prevents a dashboard from treating transport success as product success.

Create a small record around the model call

The example below is a proposed application event, using invented values and custom field names. It does not claim conformance to a telemetry transport schema.

{
  "event_name": "model_call.client_finished",
  "workflow_id": "workflow-example",
  "model_call_id": "call-example",
  "operation": "draft_answer",
  "requested_model": "example-model",
  "observed_model": null,
  "client_outcome": "completed",
  "client_duration_ms": 1240,
  "usage_status": "unavailable",
  "configuration_revision": "support-v3",
  "content_recorded": false
}

The record distinguishes the requested model from a model identity returned by the service, if one is available. It also makes missing usage visible. An empty response field should not force an invented value. Document which fields come from application configuration, which come from provider responses, and which are calculated by your instrumentation.

Measure timings with named boundaries

Choose one definition for the total workflow duration, then define narrower measurements only when they answer a question. Queue waiting, retrieval, model response, tool execution, and answer delivery can be useful boundaries. Keep units visible in field names or schema documentation. Avoid a generic duration field that means different things in different event families.

For streamed output, distinguish first observable response content from the end of the stream. State which callback starts and stops each timer. If the client disconnects before the final response, record cancellation or interruption at the boundary you observed. Do not quietly label that partial observation as a fully completed generation.

Examine a single request timeline before comparing aggregate values. Parallel operations overlap, so adding their durations can exceed the user’s actual waiting time. A trace can show that overlap while logs explain individual outcomes. Our comparison of logs, metrics, and traces helps choose the signal for each part of that investigation.

Make retries and tool calls understandable

Give the user-facing workflow an identifier and give each logical model call its own identifier. If your code observes individual network attempts, record their sequence and outcomes separately. Keep one final outcome for the logical operation. When retries occur inside a library and are not exposed, record that limitation rather than guessing how many happened.

Tool execution needs similar boundaries. Record the tool’s stable name, whether it started, and the outcome actually returned to the application. A model requesting a tool is one event; the application executing it is another. For an operation with an external effect, an ambiguous response should remain visibly ambiguous until the application can reconcile the result.

Choose one instrumentation layer to own each event family. A gateway, client wrapper, and framework may all observe related work. Assigning ownership avoids treating three descriptions of the same call as three independent calls. Preserve the relationship between records when multiple perspectives are useful.

Record configuration that explains behavior

A configuration revision can make a change understandable without copying the full instructions into every log. Keep revisions for the prompt template, retrieval configuration, tool set, and output validator in the project’s normal change history. Log the relevant revision identifiers with the workflow so an investigation can inspect the correct configuration under its existing access controls.

Record only settings your application actually supplied or observed. When a gateway selects a backend model, retain the distinction between your requested route and any returned model identity. If the backend identity is hidden, explain that in the contract. The purpose is to preserve evidence, not make the record look more complete than it is.

Decide deliberately whether to capture content

The OpenTelemetry conventions for generative AI spans describe logical operation boundaries and identify model instructions, inputs, and outputs as potentially sensitive content that instrumentation should not capture by default. The document is marked Development. When adopting these conventions, pin a reviewed revision and verify the behavior of your chosen instrumentation.

For an initial implementation, build an explicit operational field list. Review exception messages and tool arguments as carefully as the main request path. A wrapper that avoids recording prompts can still expose text by serializing an entire exception or tool result. Prefer synthetic cases when debugging instrumentation behavior.

If content is necessary for a particular evaluation, define who can access it, why it is collected, and when it should be removed. Keep that decision visible alongside the instrumentation change. The log redaction and retention guide provides a practical framework for reviewing those controls.

Keep usage and answer evaluation interpretable

Usage metadata can help explain operational changes, but it needs provenance. Label whether a count came from the provider, a local estimate, or a later reconciliation. Keep missing counts separate from known zero counts in your application records. The Token Logger guide develops this accounting design without collecting credentials or message content.

For answer evaluation, define the task and rubric before creating a score. A format validator might check required fields, while a human review might assess whether a response answers the question. Store the evaluation method and revision with its result. Do not merge different rubrics into one average simply because they produce numbers with the same range.

Test the paths an incident will expose

Run a controlled set of cases: normal completion, provider rejection, timeout, client cancellation, malformed output, and tool failure. Check the final stored records as well as the application’s local output. For each case, ask whether an operator can identify the failed boundary without reading private content.

Add a retry case and a duplicate delivery case. Confirm that repeated delivery does not inflate the count of logical operations. Check a missing usage response too, because a successful answer can still leave incomplete accounting. These exercises turn instrumentation from a collection of fields into an investigation tool.

Review instrumentation changes as behavior changes

Repeat a small synthetic workflow when upgrading a client library or adding a framework integration. Compare the stored events before and after the change. Look for renamed attributes, duplicate operations, changed timing boundaries, and unexpected content fields. A dashboard can continue rendering while its underlying meaning changes.

Keep one person or team responsible for the event contract. That owner should be able to explain which layer records each operation and how a configuration update affects historical comparisons. This is especially useful when several teams share a gateway but maintain different application wrappers.

Good AI logging makes observed behavior explainable. Start with workflow boundaries, precise outcomes, configuration revisions, and explicit gaps. Expand the record when a real operational question requires more evidence, and preserve the difference between application observations and conclusions about answer quality.