Skip to content

Handling Signals as PID 1 in Containers

Containers stop processes by sending SIGTERM, waiting a grace period, then sending SIGKILL. An asyncio service that handles SIGTERM can drain requests and close connections in that window; one that does not is killed mid-flight. Containers add a twist that surprises even experienced developers: the process running as PID 1 inside its PID namespace does not get the default action for signals it has not installed a handler for — the kernel simply ignores them. Measured with Docker on Linux and a python:3.12-slim image running an asyncio program that slept with a finally block, using docker stop with a 10-second grace period: with Python as PID 1 and no handler, the stop took 10.26 s and ended in SIGKILL (exit 137) — the SIGTERM was ignored and the cleanup never ran. With a SIGTERM handler that cancelled the main task, the stop took 0.44 s, exit 0, cleanup printed. With docker run --init and no handler, 0.22 s but exit 143 and no cleanup. Starting Python through sh -c "python ..." took 10.25 s and exit 137 even with a handler, because the shell, not Python, was PID 1. This guide makes container stops fast and clean.

Prerequisites

1. See what PID 1 does with SIGTERM

A program with no signal handler, run as the container's main process:

import asyncio


async def main() -> None:
    try:
        await asyncio.sleep(3600)
    finally:
        print("cleanup ran", flush=True)

asyncio.run(main())
docker run -d --name app python:3.12-slim python app.py
time docker stop -t 10 app        # 10.26 s; exit code 137; no "cleanup ran"

Outside a container, SIGTERM with no handler terminates the process immediately. As PID 1 it was ignored: Docker waited the full grace period and then sent SIGKILL, which cannot be caught, so nothing in the finally block ran. In Kubernetes the same happens with terminationGracePeriodSeconds (30 seconds by default): every rollout waits the full period per pod, and every pod is killed rather than stopped.

Verify: docker stop on your container returns in well under the grace period, and the container's exit code is 0 (or your chosen code), not 137.

docker stop -t 10 against an asyncio program A grid of 6 rows by 4 columns. docker stop -t 10 against an asyncio program how Python was started stop took exit cleanup ran python as PID 1, no handler 10.26 s 137 (SIGKILL) no python as PID 1, SIGTERM handler 0.44 s 0 yes --init (tini), no handler 0.22 s 143 no --init (tini), SIGTERM handler 0.41 s 0 yes sh -c "python ...", handler 10.25 s 137 no sh -c "exec python ...", handler 0.44 s 0 yes python:3.12-slim on Docker, Linux; 0.2 s of the handler cases is a simulated drain.

2. Install a SIGTERM handler that cancels the main task

The handler turns SIGTERM into a cancellation of the main task, so every finally and async with in the program runs:

import asyncio
import signal
import sys


async def main() -> int:
    loop = asyncio.get_running_loop()
    loop.add_signal_handler(signal.SIGTERM, asyncio.current_task().cancel)
    try:
        await serve_forever()
    except asyncio.CancelledError:
        log.info("SIGTERM: draining")
        await drain(timeout=20)                  # finish in-flight work within the grace period
        return 0
    finally:
        await close_resources()

if __name__ == "__main__":
    sys.exit(asyncio.run(main()))

Measured: 0.44 s from docker stop to exit, exit code 0, cleanup output present — the 0.2 s simulated drain plus process exit. A handler is what makes the difference for PID 1, because the kernel's "ignore unhandled signals" rule does not apply to signals with handlers. The drain budget must fit inside the grace period with margin, as discussed in shutting down asyncio pods in Kubernetes. SIGINT is handled by asyncio.run already; SIGTERM is not.

Verify: logs from a docker stop show the drain message and the cleanup, and the exit code is 0.

3. Use exec form, or exec in shell scripts

Dockerfile CMD and ENTRYPOINT have two forms. The shell form wraps the command in /bin/sh -c, making the shell PID 1:

# Shell form: sh is PID 1 and does not forward SIGTERM to python
CMD python -m app

# Exec form: python is PID 1 and receives SIGTERM
CMD ["python", "-m", "app"]
#!/bin/sh
# entrypoint.sh: do setup, then replace the shell with python
set -e
python -m app.migrate
exec python -m app            # exec: python takes over PID 1 and its signals

Measured: with sh -c "python handler.py" the stop took 10.25 s and ended in SIGKILL, even though Python had a handler — the signal went to the shell, which, as PID 1 with no handler, ignored it. With sh -c "exec python handler.py" it took 0.44 s and exited 0. Entrypoint scripts that do setup before starting the application are common; the exec on the last line is what makes them signal-safe.

Verify: docker exec <container> cat /proc/1/cmdline | tr '\0' ' ' shows the Python command, not a shell.

Where SIGTERM goes with a shell-form command A sequence of 5 messages between 3 participants. Where SIGTERM goes with a shell-form command docker stop sh (PID 1) python (PID 7) SIGTERM PID 1, no handler: ignored wait 10 s grace period SIGKILL (whole container) killed: no cleanup, exit 137 Measured: 10.25 s and exit 137 despite python's own SIGTERM handler.

4. Add an init when Python spawns children

docker run --init (or tini as the entrypoint) runs a tiny init as PID 1 that forwards signals to its child and reaps zombies:

FROM python:3.13-slim
RUN apt-get update && apt-get install -y --no-install-recommends tini && rm -rf /var/lib/apt/lists/*
ENTRYPOINT ["/usr/bin/tini", "--"]
CMD ["python", "-m", "app"]

Measured: with --init and no handler, the stop took 0.22 s — tini forwarded SIGTERM and Python, no longer PID 1, took the default action — but the exit code was 143 and the cleanup did not run, because the default action is immediate termination. An init solves forwarding and zombie reaping, which matters when the application spawns subprocesses or browsers, as in cleaning up Playwright on cancellation; it does not replace the handler. With both, the stop took 0.41 s and exited 0 with cleanup. In Kubernetes, shareProcessNamespace or an init in the image does the same job.

Verify: ps inside the container shows tini as PID 1 and no <defunct> processes after child processes exit.

5. Test the stop path in CI

Signal handling regresses silently — an entrypoint script loses its exec, a base image changes its CMD — so test the actual container:

#!/bin/sh
# ci/stop_test.sh
set -e
docker run -d --name stop-test "$IMAGE"
sleep 3                                        # let it start serving
start=$(date +%s)
docker stop -t 10 stop-test
elapsed=$(( $(date +%s) - start ))
code=$(docker inspect -f '{{.State.ExitCode}}' stop-test)
docker logs stop-test | grep -q "cleanup ran"
docker rm stop-test
[ "$elapsed" -lt 8 ] && [ "$code" -eq 0 ]

The test asserts three things the measurements showed can go wrong independently: the stop is fast (the signal arrived), the exit code is clean (the process exited itself rather than being killed), and the cleanup log line is present (the handler ran the shutdown path). The equivalent for Kubernetes is a rollout in a test cluster with a check for OOMKilled or Error terminations, as in Graceful Shutdown & Signal Handling.

Verify: the CI job fails when the exec is removed from the entrypoint or the handler is deleted.

What does this container need to stop cleanly? A decision on How is the process started with 4 outcomes. What does this container need to stop cleanly? How is the process started? python directly, exec form SIGTERM handler 0.44 s, exit 0 through sh or an entrypoint script exec form, or exec python else 10.25 s, 137 spawns children (subprocess, browser) tini or --init, plus handler reaping + forwarding any image CI stop test: time, exit code, cleanup log catch regressions As PID 1, an unhandled SIGTERM is ignored, not fatal.

Verification

Container stops are clean when:

  • A SIGTERM handler cancels the main task, and the program drains within the grace period.
  • The command uses exec form, or entrypoint scripts end with exec.
  • An init process runs as PID 1 when the application spawns children.
  • A CI test checks stop time, exit code and the cleanup log line.

Diagnostic Hook: chart container exit codes across deploys. Code 137 with OOMKilled=false on routine rollouts means the signal never reached a handler — a missing handler or a shell PID 1 — and every such stop also took the full grace period, which shows up as slow rollouts.

Pitfalls & edge cases

  • No handler as PID 1. Measured: SIGTERM ignored, 10.26 s, exit 137, no cleanup.
  • Shell-form CMD. Measured: 10.25 s and exit 137 even with a handler in Python.
  • Relying on an init alone. Measured: fast, but exit 143 and no cleanup.
  • Drain longer than the grace period. The handler runs, then SIGKILL arrives anyway.

Frequently Asked Questions

Why does my Python container take 10 seconds to stop?

Python is PID 1 and has no SIGTERM handler, and the kernel ignores unhandled signals for PID 1; Docker waits the grace period and sends SIGKILL. In testing that took 10.26 s and exited 137. Install a SIGTERM handler or run with an init.

How do I handle SIGTERM in an asyncio Docker container?

Call loop.add_signal_handler(signal.SIGTERM, asyncio.current_task().cancel) in the main coroutine, catch CancelledError to drain, and return an exit code. The stop took 0.44 s with exit code 0 in testing.

Does docker run --init fix signal handling?

It forwards signals and reaps zombies, so the stop is fast (0.22 s in testing), but without a handler Python is terminated immediately with exit 143 and no cleanup. Use it together with a handler.

Why doesn't my entrypoint script pass SIGTERM to Python?

The shell is PID 1 and does not forward the signal. End the script with exec python ... so Python replaces the shell; with exec the stop took 0.44 s instead of 10.25 s.