Skip to content

Timing Out Subprocesses in asyncio

Putting asyncio.wait_for around a subprocess call stops your waiting, not the subprocess. Tested on Python 3.14: after wait_for(proc.communicate(), 1.0) timed out, the child was still running with returncode None. Killing it is the obvious next step, and it has a trap of its own: a child that had written 1 MB to a stdout pipe nobody was reading was killed with SIGKILL, and await proc.wait() was still blocked three seconds later — on Python 3.12, 3.13 and 3.14 alike — because asyncio waits for the pipes to reach end-of-file and the unread data kept them open. Reading stdout to the end released it immediately; proc.communicate() after the kill returned in 2 ms. This guide writes a timeout helper that kills, reaps and keeps partial output, and escalates from SIGTERM to SIGKILL.

Prerequisites

1. Kill the process when the timeout fires

The timeout cancels the coroutine that was waiting. The process does not know and carries on:

proc = await asyncio.create_subprocess_exec(*cmd, stdout=asyncio.subprocess.PIPE)
try:
    out, _ = await asyncio.wait_for(proc.communicate(), timeout=1.0)
except TimeoutError:
    # measured: proc.returncode is None and /proc/<pid> exists - still running
    proc.kill()
    await proc.communicate()        # reap it and close the pipes (see step 2)
    raise

Measured: after the timeout, returncode was None and the process was alive; after kill() and reaping, returncode was -9 (killed by signal 9). Without the kill, every timed-out call leaves a process running for as long as it likes, and a retry loop around a hung command fills the machine with copies of it.

Verify: after a timeout in a test, proc.returncode is set and no process with that PID exists.

2. Reap with communicate(), not wait(), when pipes are attached

After a kill, await proc.wait() returns only when the process has exited and its pipes have been closed. If the child wrote more to a pipe than the stream reader buffers, and nobody reads it, the pipe never reaches end-of-file from asyncio's point of view:

proc = await asyncio.create_subprocess_exec(*cmd_that_writes_1mb, stdout=asyncio.subprocess.PIPE)
await asyncio.sleep(1.0)
proc.kill()
await proc.wait()               # measured: still blocked 3 s later on 3.12, 3.13 and 3.14

# Either read the pipes to EOF...
await proc.stdout.read()
await proc.wait()               # returned at once, returncode -9

# ...or let communicate() do both
await proc.communicate()        # measured: returned in 0.002 s after the kill

The process itself was gone (a zombie waiting to be reaped); it was the unread pipe data that kept wait() from completing. communicate() reads every attached pipe to end-of-file and then waits, so it always completes once the process has exited. Use it as the reaping call whenever stdout or stderr is a pipe, and use wait() only for processes whose output goes elsewhere (inherited, a file, or DEVNULL).

Verify: a test that kills a chatty process after a timeout completes its cleanup in milliseconds, not at the test runner's own timeout.

Reaping a killed process with 1 MB of unread stdout A grid of 3 rows by 3 columns. Reaping a killed process with 1 MB of unread stdout after SIGKILL result why await proc.wait() blocked > 3 s (3.12-3.14) pipe never reached EOF stdout.read() then wait() returned at once pipe drained await proc.communicate() returned in 0.002 s reads all pipes, then waits With pipes attached, communicate() is the safe way to reap.

3. Keep partial output by streaming it

communicate() returns output only if it completes. When it is cancelled by a timeout, whatever it had read so far is gone — tested, a second communicate() after the kill returned 0 bytes of a 1 MB output. When partial output matters (logs of a hung build, progress of a stuck job), read it as it arrives:

async def run_with_timeout(*cmd: str, timeout: float) -> tuple[int, list[bytes], bool]:
    proc = await asyncio.create_subprocess_exec(
        *cmd, stdout=asyncio.subprocess.PIPE, stderr=asyncio.subprocess.STDOUT,
    )
    lines: list[bytes] = []

    async def collect() -> None:
        async for line in proc.stdout:
            lines.append(line)

    reader = asyncio.create_task(collect())
    timed_out = False
    try:
        async with asyncio.timeout(timeout):
            await proc.wait()
    except TimeoutError:
        timed_out = True
        proc.kill()
    await reader                     # reads to EOF, which arrives once the process is dead
    await proc.wait()
    return proc.returncode, lines, timed_out

Tested with a process that printed three lines and then hung: after the 1 s timeout and kill, all three lines were kept. The reader task keeps the pipe drained the whole time, so the deadlock from step 2 cannot occur and the child never blocks on a full pipe either. Merging stderr into stdout keeps the interleaving that a human reading the log expects.

Verify: a command that prints and then hangs returns its printed lines together with timed_out=True.

A subprocess that hangs, with a streaming timeout helper 3 lanes over time. A subprocess that hangs, with a streaming timeout helper child prints 3 lines hangs killed reader task collects lines waiting EOF helper timeout(1.0) kill, reap return lines, timed_out time → Draining continuously keeps both the output and the reaping safe.

4. Escalate from SIGTERM to SIGKILL

SIGKILL loses whatever the process would have cleaned up. Ask first with SIGTERM, give it a short grace period, then kill — and bound the whole sequence:

async def stop(proc: asyncio.subprocess.Process, grace: float = 5.0) -> int:
    if proc.returncode is not None:
        return proc.returncode
    proc.terminate()                                       # SIGTERM
    try:
        await asyncio.wait_for(proc.communicate(), grace)
    except TimeoutError:
        proc.kill()                                        # SIGKILL: cannot be ignored
        await proc.communicate()
    return proc.returncode

Tested with a process that ignores SIGTERM: it was still running after the 2 s grace period, and the SIGKILL that followed ended it with return code -9. Programs that handle SIGTERM — most servers, ffmpeg, test runners — exit cleanly within the grace period and leave no partial files. For commands that spawn children, send both signals to the process group, as in killing subprocess trees on cancellation.

Verify: a SIGTERM-ignoring test process is killed after the grace period, and a well-behaved one exits with its normal shutdown code.

5. Shield cleanup and report timeouts clearly

The timeout handler itself runs during cancellation. If the surrounding task is cancelled again — a client disconnect, a shutdown — the cleanup can be interrupted halfway, leaving the process running. Shield it, and surface timeouts as their own error type:

class CommandTimeout(Exception):
    def __init__(self, cmd, timeout, output):
        super().__init__(f"{cmd[0]} timed out after {timeout}s")
        self.output = output


async def run(*cmd: str, timeout: float) -> bytes:
    proc = await asyncio.create_subprocess_exec(*cmd, stdout=asyncio.subprocess.PIPE,
                                                stderr=asyncio.subprocess.STDOUT)
    try:
        async with asyncio.timeout(timeout):
            out, _ = await proc.communicate()
            return out
    except TimeoutError:
        await asyncio.shield(stop(proc))
        raise CommandTimeout(cmd, timeout, b"") from None
    except asyncio.CancelledError:
        await asyncio.shield(stop(proc, grace=1.0))       # caller gave up: stop quickly
        raise

A dedicated exception lets callers distinguish "the command failed" (non-zero exit) from "the command did not finish" and retry, alert or degrade accordingly. Pick timeouts per command from its observed duration distribution, generously above the p99, and record them; a timeout that fires routinely is a capacity problem, not a safety net.

Verify: cancelling the caller during the command stops the process within the short grace period, and timeouts appear in logs as CommandTimeout with the command name.

How should this subprocess call be bounded? A decision on What does the command produce with 3 outcomes. How should this subprocess call be bounded? What does the command produce? small output, all or nothing timeout around communicate() kill, then communicate() logs you need on failure stream with a reader task keep partial lines spawns children process group killpg TERM then KILL Every path ends with the process dead and reaped, and the pipes drained.

Verification

Subprocess timeouts are correct when:

  • A timeout always kills the process, never just stops waiting.
  • Reaping uses communicate() or drains pipes before wait().
  • SIGTERM comes first, with SIGKILL after a bounded grace period.
  • Cleanup is shielded and timeouts surface as a distinct error.

Diagnostic Hook: count timeouts per command and the number of child processes over time. A child count that steps up after each timeout means processes are not being killed; cleanup that itself takes as long as the test or request timeout means wait() is blocked on an undrained pipe.

Pitfalls & edge cases

  • wait_for without a kill. Measured: the child kept running.
  • wait() after kill with unread pipes. Measured: blocked indefinitely on 3.12–3.14.
  • Expecting partial output from a cancelled communicate(). It is lost.
  • SIGKILL only. Programs cannot clean up temporary files or locks.

Frequently Asked Questions

Does asyncio.wait_for kill a subprocess on timeout?

No. It cancels the waiting coroutine; the process keeps running. In testing the child was still alive with returncode None after the timeout. Call proc.kill() and reap it.

Why does proc.wait() hang after killing a subprocess?

asyncio waits for the stdout and stderr pipes to close, and unread data keeps them open. In testing on Python 3.12 to 3.14, wait() blocked after SIGKILL until stdout was read; communicate() returned in 2 ms.

How do I keep a subprocess's output when it times out?

Read stdout line by line in a separate task while waiting for the process, then kill on timeout and await the reader, which receives everything printed before the kill.

Should I send SIGTERM or SIGKILL on a subprocess timeout?

SIGTERM first, so the program can clean up, then SIGKILL after a short grace period if it has not exited. A process ignoring SIGTERM was ended by the follow-up SIGKILL in testing.