Useful structured logging starts with a question: what will someone need to understand when an operation behaves unexpectedly? A line saying “request failed” establishes that something happened, but leaves the reader to reconstruct the service, operation, timing, and outcome. A structured event puts those facts into named fields with meanings that remain consistent across requests.

This guide proposes a small logging contract for an application team. It is a design starting point, not a universal schema or a claim that more logging automatically improves reliability. Begin with one important workflow, make its events understandable, and expand only when a new field answers a real question. The Log Mic overview introduces the wider collection and analysis workflow.

Choose the questions before the fields

Imagine a document conversion service. An operator needs to distinguish unsupported files, dependency timeouts, cancelled jobs, and successful conversions. A developer needs to compare behavior before and after a release. Someone planning capacity needs to understand work arriving and work completing. These questions suggest operation names, outcomes, release identifiers, and durations.

Write three investigation questions beside the code you intend to instrument. For example: which conversions failed after the latest deployment, which dependency was involved, and did a retry eventually succeed? Then sketch the query that would answer each question. If the query depends on searching a sentence fragment, consider giving that fact a field. If nobody can describe a use for a proposed field, leave it out of the first version.

Define a small event contract

A logging contract states the expected field names, types, units, and missing-value rules. Keep this document close to the code so application changes and logging changes can be reviewed together. The following is an illustrative application schema; its names are not an OpenTelemetry wire format.

FieldProposed meaning
event_nameA stable operation milestone, such as conversion.completed.
service and environmentThe application component and deployment environment.
request_idAn opaque identifier linking events from one request.
outcomeA documented value such as success, failure, or cancelled.
duration_msElapsed time for a precisely defined operation boundary.
schema_versionThe version of this application event contract.

Separate an event name from its explanation

Use the event name for grouping and a short message for reading. A name such as conversion.completed should remain the same when a filename, duration, or outcome changes. The message can explain that processing ended, while other fields carry the details. Avoid placing customer names or unique document identifiers inside the event name; those values create new groups whenever they change.

Make an event useful on its own

This invented record describes a single conversion attempt. Its numbers are illustrative, and its identifiers refer to no real request.

{
  "schema_version": 1,
  "event_name": "conversion.completed",
  "service": "document-worker",
  "environment": "test",
  "request_id": "request-example",
  "attempt": 1,
  "outcome": "failure",
  "error_class": "dependency_timeout",
  "duration_ms": 820,
  "release": "example-release"
}

Notice what the record lets a reader decide without opening a second system: the work happened in a test environment, the failure involved a dependency timeout, and it occurred on the first attempt. The duration needs a definition in the contract: here it covers conversion processing from the worker accepting the attempt until that attempt ends. It does not include time waiting in a queue.

Before shipping an event, ask a teammate to explain it without reading the emitting function. If they confuse an attempt with a completed job, or interpret a duration as total user waiting time, improve the field names or contract. This small review often exposes ambiguity earlier than a dashboard does.

Keep time and correlation explicit

The OpenTelemetry Logs Data Model distinguishes event time from observation time and defines optional trace and span context. It also separates a record body, event attributes, and information describing its source. Those distinctions are a useful reference when mapping an application schema into a common telemetry model.

In your own contract, label timestamps according to the clock and boundary they represent. An event written by a worker and collected later has two useful moments: when the worker recorded it and when collection observed it. Preserving both helps you ask whether apparent delay belongs to the application or the logging path. Do not silently replace a missing source timestamp with collection time while keeping the original label.

Choose one request identifier at the workflow boundary and pass it through the parts you control. If tracing is already installed, use its established context rather than inventing an incompatible tracing scheme. The guide to logs, metrics, and traces explains the different questions these signals can answer together.

Set rules for outcomes and severity

Outcome describes what happened to an operation; severity describes how your team intends to treat a particular event. Decide whether an expected validation rejection belongs at an informational level, and when a retried dependency failure deserves a warning. Document these decisions with examples. Otherwise two services can report equivalent situations in ways that make a shared view misleading.

A failed attempt followed by a successful retry should not leave the job outcome ambiguous. Give each attempt a number, and write a separate final job outcome when the workflow finishes. For an early rollout, prefer this small set of meaningful milestones to a line for every function call. Add detailed debugging events only around an investigation with a defined end date.

Control what enters the record

Build events from an explicit list of fields. Passing an entire request object to the logger makes the event depend on everything that object might contain, including fields added by another developer later. Prefer safe categories such as document type and processing mode when they answer the question. Keep uploaded content, authorization values, and unrestricted headers outside routine operational logs.

Validate field lengths and types before serialization. If a value is too large, preserve a clear truncation indicator rather than producing a record that appears complete. Use synthetic test strings containing line breaks, quotation marks, and unusual characters to check that one logical event remains one parseable record. Our redaction and retention guide develops the decisions around access and deletion.

Define empty and missing values

Decide whether a missing attribute means unknown, irrelevant, or intentionally removed. Those meanings can matter during an incident. A duration of zero states something different from a duration that was never measured. Preserve the distinction in the record instead of forcing every event into a shape that looks complete.

When a field changes meaning, treat that as a contract change even if its name stays the same. For example, expanding a duration to include queue waiting will alter comparisons with older events. Introduce a new field or schema version, explain the transition, and update the affected queries together. Keep old fixtures so you can confirm that records from both versions remain interpretable.

Roll out with a query and a failure exercise

Start with a single service and save a few representative fixtures: success, expected rejection, timeout, cancellation, and retry success. Validate required fields, numeric durations, accepted outcomes, and the absence of disallowed fields. These checks should express the event contract rather than repeat every detail of the logging implementation.

Next, send those records through the actual collection path used by the service. Read the stored result, because a correct application record can still be renamed, truncated, or dropped during processing. Run the investigation queries drafted at the beginning. Confirm that a release filter selects the intended events and that missing information remains visibly missing.

Exercise a temporarily unavailable destination in a controlled environment. Decide how the application should behave when logs cannot be delivered, how much buffering is acceptable, and how loss will become visible. Record that decision. A logging plan is incomplete when it only describes the successful delivery path.

Improve the contract through real investigations

Review the events after an incident or support investigation. Which field answered the question? Which field was missing? Which large payload was never opened? Use those observations to refine the schema and remove unnecessary detail. Keep a short change history so a renamed field does not quietly break an old query.

Structured logging becomes useful when its meaning survives changes in code, teams, and storage. Start with clear operation boundaries, consistent fields, and a few tested queries. That gives every additional event a purpose and gives the next reader a fair chance of understanding what actually happened.