Skip to content

Killing Subprocess Trees on Cancellation

Cancelling the asyncio task that awaits a subprocess does not stop the subprocess. And stopping the subprocess does not stop its children. Tested on Linux with Python 3.14, running sh -c "sleep & sleep & wait" — a shell with two child processes — and cancelling the awaiting task: with no cleanup, the shell and both children kept running; with proc.kill() in the cancellation handler, the shell died but both children survived, re-parented to init; starting the shell in its own process group and killing the group with os.killpg() left zero survivors. Leaked processes accumulate across requests or retries until something runs out — CPU, memory, file locks, or ports. This guide makes subprocess lifetimes follow task lifetimes, with a graceful stop before the hard one.

Prerequisites

1. See what cancellation leaves behind

The awaiting task and the process are separate things. Cancelling one tells the other nothing:

proc = await asyncio.create_subprocess_exec("sh", "-c", "worker-a & worker-b & wait")

task = asyncio.create_task(proc.wait())
task.cancel()            # the task ends; sh, worker-a and worker-b keep running

# Even killing the direct child is not enough:
proc.kill()              # sh dies; worker-a and worker-b are re-parented and keep running

Measured: no cleanup left the shell and 2 children running; proc.kill() left the 2 children. Shells, build tools, browsers, npm, ffmpeg wrappers and many CLIs spawn children, so "the process I started" is rarely the whole tree. The orphans are hard to notice because nothing in your program refers to them any more.

Verify: after cancelling a task in a test, pgrep -f for the command's children returns nothing.

Child processes still running after cancelling the task 3 horizontal bars comparing no cleanup with the others. Child processes still running after cancelling the task no cleanup 2 (and the shell) proc.kill() 2, re-parented start_new_session + os.killpg() 0 Python 3.14 on Linux; command sh -c 'sleep 77.1 & sleep 77.2 & wait'. Killing the process you started is not killing what it started.

2. Start the subprocess in its own process group

A process group collects a process and every descendant that does not deliberately leave it. Start the subprocess as the leader of a new group, then signal the whole group:

import os
import signal


async def start(*cmd: str) -> asyncio.subprocess.Process:
    return await asyncio.create_subprocess_exec(
        *cmd,
        start_new_session=True,            # new session and process group; pgid == proc.pid
        stdout=asyncio.subprocess.PIPE,
        stderr=asyncio.subprocess.PIPE,
    )


def kill_tree(proc: asyncio.subprocess.Process, sig: int = signal.SIGKILL) -> None:
    try:
        os.killpg(proc.pid, sig)           # every process in the group
    except ProcessLookupError:
        pass                               # already gone

start_new_session=True also detaches the child from your terminal, so a Ctrl-C in the terminal reaches your program but not the subprocess directly — your program decides. process_group=0 (Python 3.11+) creates a new group without a new session, if you need the terminal relationship kept. Measured, os.killpg after start_new_session left no survivors.

Verify: ps -o pid,pgid,cmd shows the subprocess and its children sharing one PGID equal to the subprocess's PID.

3. Tie the tree to the task with a context manager

Wrap start and cleanup so every exit path — success, exception, cancellation — kills the tree:

from contextlib import asynccontextmanager


@asynccontextmanager
async def subprocess_tree(*cmd: str, grace: float = 5.0):
    proc = await start(*cmd)
    try:
        yield proc
    finally:
        if proc.returncode is None:
            kill_tree(proc, signal.SIGTERM)                    # ask nicely first
            try:
                await asyncio.shield(asyncio.wait_for(proc.communicate(), grace))
            except (TimeoutError, asyncio.CancelledError):
                kill_tree(proc, signal.SIGKILL)                # then insist
                await asyncio.shield(proc.communicate())
        else:
            kill_tree(proc, signal.SIGKILL)                    # leader exited; children may not have


async def transcode(src: str, dst: str) -> None:
    async with subprocess_tree("ffmpeg-wrapper.sh", src, dst) as proc:
        await proc.communicate()                               # drains the pipes while waiting

The shield keeps the cleanup's own awaits from being interrupted by a second cancellation — without it, cancelling twice could skip the SIGKILL. Killing the group even after the leader has exited catches children that outlived it. Reaping with proc.communicate() rather than proc.wait() matters because the pipes are attached: after a kill, wait() can block until unread output is drained, as shown in timing out subprocesses in asyncio.

Verify: cancel transcode at random points in a test loop; no process from the group remains afterwards.

Stopping a subprocess tree A flow of 5 stages. Stopping a subprocess tree start_new_session=True own process group task ends or is cancelled finally runs killpg(SIGTERM) graceful stop wait up to grace shielded killpg(SIGKILL) then reap Graceful first, certain second, and always for the whole group.

4. Give the tree a graceful stop

SIGKILL cannot be caught; it leaves temporary files, partial outputs and held locks behind. Send SIGTERM first so well-behaved programs can clean up, and make the grace period long enough for them to do so:

GRACE_BY_TOOL = {
    "ffmpeg": 2.0,          # finalizes the container on SIGTERM
    "pg_dump": 5.0,
    "pytest": 10.0,         # runs fixture teardown
}


async def stop(proc, tool: str) -> int:
    kill_tree(proc, signal.SIGTERM)
    try:
        await asyncio.wait_for(proc.communicate(), GRACE_BY_TOOL.get(tool, 5.0))
    except TimeoutError:
        kill_tree(proc, signal.SIGKILL)
        await proc.communicate()
    return proc.returncode

Some programs prefer SIGINT (interactive tools treat it as Ctrl-C) or have their own stop command; check what the tool documents. Shell scripts deserve attention: a shell running a foreground command may not exit on SIGTERM until that command does, which is another reason to signal the whole group rather than only the shell. The return code tells you how it ended: negative values are the signal number that killed it.

Verify: a test that cancels mid-run finds the tool's own cleanup done (no leftover temp files) in the common case, and no processes in any case.

5. Add a backstop for when your process dies

Process groups only help if your program gets to run its cleanup. If it is killed with SIGKILL or crashes, nobody signals the group. Two backstops on Linux:

import ctypes
import signal

PR_SET_PDEATHSIG = 1
libc = ctypes.CDLL("libc.so.6", use_errno=True)


def die_with_parent() -> None:
    libc.prctl(PR_SET_PDEATHSIG, signal.SIGKILL)      # runs in the child before exec


proc = await asyncio.create_subprocess_exec(*cmd, start_new_session=True, preexec_fn=die_with_parent)

PR_SET_PDEATHSIG makes the kernel send the child a signal when its parent thread exits — it covers the direct child, not grandchildren, and preexec_fn is not safe in multi-threaded programs, so prefer the second option where you can: run the service under systemd or in a container, where the cgroup contains every descendant and stopping the unit or container kills all of them. In Kubernetes, the pod's cgroup does this when the container stops. Treat in-process cleanup as the normal path and the cgroup as the guarantee.

Verify: kill -9 your service during a run; under systemd (KillMode=control-group) or in a container, no descendant survives the stop.

How should this subprocess be contained? A decision on Can the command spawn children with 3 outcomes. How should this subprocess be contained? Can the command spawn children? no, single binary proc.terminate()/kill() still reap it yes, or unsure (shells, tools) new process group + killpg graceful then hard long-running service plus cgroup backstop systemd, container Assume a tree unless you know it is a single process.

Verification

Subprocess trees are contained when:

  • Every subprocess starts in its own process group.
  • Every exit path kills the group, with SIGTERM, a grace period, then SIGKILL.
  • Cleanup is shielded from repeated cancellation and reaps the leader.
  • A cgroup backstop covers crashes of the parent.

Diagnostic Hook: periodically count processes whose parent is PID 1 (or your container's init) and whose command matches tools you launch. A growing count means trees are escaping — usually a code path that kills the leader only, or a finally that is cancelled before it completes.

Pitfalls & edge cases

  • Relying on cancellation to stop the process. Measured: shell and children kept running.
  • proc.kill() on a shell. Measured: 2 grandchildren survived.
  • Unshielded cleanup. A second cancellation can skip the SIGKILL.
  • No backstop. A crashed parent leaves the whole tree behind.

Frequently Asked Questions

Does cancelling an asyncio task kill its subprocess?

No. The task ends, but the process keeps running. In testing, a shell and its two children all kept running after the awaiting task was cancelled.

Why do child processes survive proc.kill() in asyncio?

proc.kill() signals only the direct child. Its own children are re-parented and continue. Start the subprocess with start_new_session=True and kill the group with os.killpg(proc.pid, sig).

How do I stop a subprocess gracefully from asyncio?

Send SIGTERM to its process group, wait for it with a timeout, and send SIGKILL to the group if it has not exited, reaping it with await proc.communicate() when its output is piped.

How do I make sure subprocesses die if my Python program crashes?

Run the program under systemd or in a container so the cgroup is cleaned up on stop; on Linux, PR_SET_PDEATHSIG can also signal a direct child when its parent dies.