Skip to content

Flushing Telemetry and Logs Before Exit

Logs, traces and metrics are buffered so they do not slow the work down — and buffers are lost when a process exits without draining them. The last seconds before a shutdown are often exactly what an incident review needs. Measured with an asyncio service writing a log line and a span per request through a QueueListener and an OpenTelemetry BatchSpanProcessor, about 1,700 requests in two seconds before SIGTERM: with no SIGTERM handler, the process died with exit code 143, leaving 924 log lines and 1,536 spans — three full export batches, the rest gone. With a handler that exited cleanly but did not flush, all 1,679 spans arrived (the SDK's exit hook flushed them) but only 911 log lines, because the queue listener's thread was never stopped. Stopping the listener and flushing the tracer provider explicitly delivered 1,699 of 1,699 of both. This guide makes the final telemetry reliable.

Prerequisites

1. Handle SIGTERM so exit hooks can run

Without a handler, SIGTERM kills the Python process immediately: no finally blocks, no atexit hooks, no buffer flushes. The first requirement is to turn SIGTERM into a normal exit:

import signal


async def main() -> None:
    stop = asyncio.Event()
    asyncio.get_running_loop().add_signal_handler(signal.SIGTERM, stop.set)
    try:
        await serve_until(stop)
    finally:
        await flush_telemetry()          # step 4


asyncio.run(main())                      # returns normally; atexit hooks then run

Measured without a handler: exit code 143 (128 + SIGTERM), 924 log lines and 1,536 spans out of about 1,700 — whatever had been exported before the kill. The OpenTelemetry SDK registers an exit hook that shuts the provider down and flushes spans, but exit hooks only run on a normal interpreter exit. That is why simply installing a handler already rescued every span in the second run.

Verify: the process exits with code 0 after SIGTERM, and its last log line is the shutdown message.

Telemetry delivered after SIGTERM, about 1,700 produced 5 horizontal bars comparing no handler: log lines with the others. Telemetry delivered after SIGTERM, about 1,700 produced no handler: log lines 924 no handler: spans 1,536 handler, no flush: log lines 911 handler, no flush: spans 1,679 handler + explicit flush: both 1,699 each Python 3.14, QueueListener with a slow file handler, OpenTelemetry BatchSpanProcessor (5 s delay, 512 per batch). Exit hooks rescued spans; only an explicit stop rescued queued logs.

2. Stop queue-based log handlers explicitly

Async services often log through a QueueHandler so slow handlers — network shippers, files on slow disks — never block the event loop. The QueueListener drains the queue in a thread, and nothing stops it automatically:

import logging.handlers
import queue

log_queue: queue.Queue = queue.Queue(-1)
listener = logging.handlers.QueueListener(log_queue, shipping_handler, respect_handler_level=True)


def configure_logging() -> None:
    root = logging.getLogger()
    root.addHandler(logging.handlers.QueueHandler(log_queue))
    listener.start()


def stop_logging() -> None:
    listener.stop()                  # processes everything queued, then joins the thread
    shipping_handler.close()

Measured: with a clean exit but no listener.stop(), only 911 of about 1,700 lines reached the file; with stop(), all 1,699. QueueListener.stop() enqueues a sentinel and waits for the thread to process everything before it, so it can take as long as the backlog needs — bound it by keeping the backlog small (step 5). From Python 3.12, logging.config.dictConfig can create the listener for you, but you still have to stop it.

Verify: the final log line written before stop_logging() appears at the log destination after every restart.

3. Force-flush span and metric exporters

Batch processors and periodic metric readers hold data for seconds by design. Flush them as the last step of shutdown, with a timeout:

from opentelemetry import metrics, trace


def flush_otel(timeout_ms: int = 5000) -> None:
    tracer_provider = trace.get_tracer_provider()
    if hasattr(tracer_provider, "force_flush"):
        tracer_provider.force_flush(timeout_ms)
        tracer_provider.shutdown()
    meter_provider = metrics.get_meter_provider()
    if hasattr(meter_provider, "force_flush"):
        meter_provider.force_flush(timeout_ms)
        meter_provider.shutdown()

Measured: with an explicit flush, 1,699 of 1,699 spans arrived. Relying on the SDK's exit hook also worked once SIGTERM was handled, but an explicit flush in the shutdown sequence makes the order deterministic — spans are flushed after the work that produces them has stopped, and before the process starts tearing down the network clients the exporter needs. The flush calls block, so call them via asyncio.to_thread if anything else on the loop must keep running.

Verify: traces for requests completed in the last second before a shutdown are visible in the tracing backend.

The order of a telemetry-safe shutdown A flow of 5 stages. The order of a telemetry-safe shutdown SIGTERM handled normal exit path drain work last spans produced log shutdown summary while logging works flush traces + metrics bounded timeout stop log listener drain the queue, exit 0 Flush telemetry after the work stops and before the tools that carry it are closed.

4. Put the flush last, within the grace period

The telemetry flush belongs at the very end of the shutdown sequence, and it needs time in the budget:

async def shutdown_sequence(server, workers, budget: float = 20.0) -> None:
    started = time.monotonic()
    await server.stop_accepting()
    await drain(workers, timeout=budget - 5.0)              # leave 5 s for the rest
    log.info("shutdown: drained in %.1fs", time.monotonic() - started)
    await close_pools()
    await asyncio.to_thread(flush_otel, 3000)               # spans and metrics
    stop_logging()                                          # logs last: captures everything above

Logging is stopped last so that everything before it — including messages about the flush itself — is captured. Give the flush a timeout shorter than what remains of the grace period; an exporter stuck on an unreachable collector must not hold the pod until SIGKILL, which would lose the logs that come after it.

Verify: shutdown logs include the "drained" line and appear at the destination, and the measured shutdown duration includes a bounded flush step.

5. Keep buffers small enough to drain

Everything buffered at the moment of shutdown has to be drained in the time available. Size buffers for that, not just for throughput:

from opentelemetry.sdk.trace.export import BatchSpanProcessor

processor = BatchSpanProcessor(
    exporter,
    max_queue_size=2048,             # bounded: excess spans are dropped (counted), not hoarded
    schedule_delay_millis=1000,      # export every second, not every five
    max_export_batch_size=512,
    export_timeout_millis=3000,
)

A shorter schedule delay means less data at risk at any moment; in the measurement, a 5-second delay left everything since the last full batch exposed. For logs, a bounded queue that drops (and counts) messages when the destination is down is better than an unbounded one that grows during an outage and then cannot be drained before exit. Ship logs to stdout and let the platform collect them, where possible — the container runtime's log driver outlives your process.

Verify: at peak load, the span queue and log queue depths stay far below what can be drained within the shutdown budget.

What does each telemetry path need at shutdown? A decision on How does this telemetry leave the process with 4 outcomes. What does each telemetry path need at shutdown? How does this telemetry leave the process? anything buffered handle SIGTERM otherwise lost at kill QueueHandler + listener listener.stop() 911 -> 1,699 lines OTel batch exporters force_flush + shutdown bounded timeout stdout logging platform collects nothing to flush in-process Unflushed buffers are where the last seconds of every incident go missing.

Verification

Telemetry survives shutdown when:

  • SIGTERM leads to a normal exit, so finally blocks and exit hooks run.
  • Queue listeners are stopped and span and metric providers flushed, explicitly.
  • The flush comes last and is bounded within the grace period.
  • Buffers are small enough to drain in the time available.

Diagnostic Hook: emit a "shutdown complete" log line and span as the very last action, and alert on deploys where it never arrives. A missing final line means the process was killed before flushing — check the exit code (143 or 137) and the shutdown duration against the grace period.

Pitfalls & edge cases

  • No SIGTERM handler. Measured: exit code 143, spans and logs lost.
  • Queue listener never stopped. Measured: 911 of about 1,700 log lines arrived.
  • Unbounded flush. A stuck exporter holds the process until SIGKILL.
  • Long batch delays. More data at risk at every moment.

Frequently Asked Questions

Why are my last logs missing after a container stops?

Either the process was killed by SIGTERM without a handler, or a queue-based log listener was never stopped. In testing, an unstopped QueueListener delivered 911 of about 1,700 lines; listener.stop() delivered all of them.

Does OpenTelemetry flush spans when a Python process exits?

Its exit hook flushes on a normal interpreter exit, which requires handling SIGTERM. With no handler, only spans already exported in full batches survived in testing.

When should telemetry be flushed during shutdown?

Last: after work has stopped and the final log lines are written, with a timeout shorter than the remaining grace period, then stop the log listener.

How do I avoid losing logs on SIGKILL?

You cannot flush on SIGKILL. Keep buffers small, export frequently, log to stdout so the platform collects it, and size the grace period so SIGKILL is rare.