A JSON logfile is easy to create and surprisingly easy to collect incorrectly. An application writes a record, a collector reads it, and a rotation job eventually replaces the file. Each component can appear healthy while a few records disappear between them. The difficult part is agreeing on what happens at boundaries: a partial write, a renamed file, an interrupted upload, or a restart during a busy minute.
Start with a workflow you can explain and test. This guide uses a hypothetical document conversion worker that writes completion and failure events to a local file. The same questions apply to other services. For an overview of the responsibilities involved, visit the Logfile Logger collection guide.
Define the record before choosing the rotation policy
Use one complete JSON object per physical line for this workflow. Treat the newline as the record boundary and escape line breaks inside string values through a JSON serializer. Avoid assembling JSON by joining untrusted strings. Pretty printed objects are convenient on a terminal, but they require a different collection contract because a single event spans several lines.
Choose a small envelope: event identifier, event timestamp, service name, severity, event name, and schema version. Add fields that describe the operation without copying its entire input. The conversion worker might record a document format, an outcome, and elapsed milliseconds. A request identifier can connect its events without putting document contents into the logfile.
Write down maximum event size, handling of invalid values, and whether the final newline is required. Decide how the producer reports a serialization failure. A silently skipped event is difficult to distinguish from an operation that never happened. The structured logging guide develops the schema decisions that make later collection predictable.
Give the writer and reader separate responsibilities
The application owns event meaning and serialization. The collector owns reading progress, transport, and acknowledgments from its destination. Rotation owns the lifecycle of local files. Avoid configurations where two independent components both rotate the same path; assign one component clear ownership of that action.
For each file, ask how the collector recognizes identity after a rename, how it notices truncation, and where it stores its reading position. These behaviors depend on the collector and filesystem. Record the actual configuration and verify the behavior with a test. A filename alone is insufficient as a design explanation when that name is repeatedly reused.
Also separate reading a line from successfully delivering it. Find out when progress is committed and what survives a crash. If a batch is accepted remotely but the local checkpoint is not saved, replay may produce duplicates. If progress is saved before durable delivery, interruption may lose records. Decide which outcomes your design permits and how you will detect them.
Choose a rotation method the application can support
The upstream logrotate manual documents an important tradeoff: copytruncate copies a file and then empties the original, with a window in which written data may be lost. It also explains that delaycompress postpones compression until the next rotation cycle and that debug mode makes no changes to logs or the state file. Rotation criteria are evaluated when logrotate runs; a size setting alone does not create continuous monitoring.
For the example worker, prefer a design where the writer can close its old handle and open the new file after rotation. Confirm the application's supported mechanism before configuring it. Do not guess a signal or reuse a command intended for another service. Test what happens if reopening fails, including ownership and permissions on the new path.
If the writer cannot reopen, document that limitation before selecting an alternative. A delayed compression policy can buy time for a reader, but it is not proof that the reader finished. Establish a handoff rule based on observed collection progress and the actual capabilities of your chosen tools.
Plan for a destination that stops accepting records
Collection needs an explicit outage policy. Suppose the worker emits 2 megabytes of serialized events per minute. A hypothetical 45 minute transport interruption creates 90 megabytes of new events before allowing for bursts, indexes, checkpoints, or other overhead. That arithmetic describes a starting buffer requirement, not a guarantee that a 90 megabyte disk allocation is sufficient.
Decide where records wait during the interruption. Keep local retention long enough for the collector to recover under the expected workload, and reserve additional space for the active file and rotation operations. Set separate observations for free space, backlog age, failed delivery attempts, and newly dropped events. A healthy process count does not answer these questions.
Choose the response to a full buffer deliberately. Blocking an application, dropping selected diagnostic events, and failing an operation have different consequences. The right choice depends on the event's purpose. Record it in the runbook so an operator does not improvise by deleting the oldest files while an incident is still being investigated.
Test rotation with a countable sequence
Build a small, observable rehearsal
In a staging environment, generate events with a unique run identifier and sequence numbers from one through ten thousand. Keep an independent expected count. Include a few deliberately large events and strings containing escaped newlines. Rotate during the run, continue writing, and wait until the destination reports that the backlog is clear.
Compare received identifiers with the expected set. Measure missing events, duplicate events, parse failures, and unexpected sequence gaps separately. A total of ten thousand received rows can conceal one missing record and one duplicate. Check both the set of identifiers and the number of occurrences of each identifier.
Change one failure at a time
Repeat the exercise with the collector stopped during rotation. Then test a collector restart after a batch upload, a temporary destination rejection, and an application restart. Finally, combine the two failures most plausible in your environment. Save the test conditions and observed results beside the configuration version.
Inspect the last line of the old file and the first line of the new file. Confirm that incomplete trailing data is handled according to your contract and does not become a valid but misleading event. These boundary checks explain failures that a dashboard of average ingestion rate can hide.
Preserve time without pretending it provides perfect order
Store when the producer says the event occurred separately from when collection observed it. Name the fields clearly and use an explicit timezone. For elapsed operation duration, record a duration measured by the application rather than asking an analyst to infer it from neighboring log timestamps.
In the conversion example, two workers may finish separate documents at nearly the same moment. Sorting by wall clock time does not establish a causal relationship between them. A run identifier, operation identifier, and sequence within an operation provide more useful context. Keep timestamp precision consistent, but do not add precision the producer did not measure.
When replaying a backlog, retain the original event time. Mark replay or collection time separately if needed. Otherwise a transport recovery can look like a sudden burst of new application failures. Confirm that searches can distinguish recent collection activity from the period in which the underlying work happened.
Make local cleanup a verified operation
A rotation count describes how many generations you retain, while your operational requirement may be expressed in hours or days. Translate between them using measured rotation frequency. A burst that triggers extra rotations can shorten the history represented by a fixed number of files.
Review active logs, rotated copies, compressed archives, rejected records, and exported incident bundles together. The redaction and retention workflow covers how to assign each copy a purpose and an expiry rule. Include a restore rehearsal in the operational plan: identify a historical run, retrieve the permitted records, and confirm that the parser still understands its schema version.
Conclusion: prove the handoffs
A reliable JSON logfile workflow is a set of tested agreements. Define complete records, give rotation one owner, verify reopening and reading progress, and rehearse outages with countable identifiers. Keep the resulting evidence close to the configuration. When a collector or application changes, repeat the relevant boundary test. That gives the next operator a concrete answer to the most useful logging question: which events reached their destination, and how do we know?



