Skip to content

Limiting Concurrent Subprocesses

asyncio.gather over create_subprocess_exec calls starts every process at once. For a handful that is fine; for a batch it fails in two ways. Each subprocess with captured output costs several file descriptors, so a large batch runs into the descriptor limit: tested with the common limit of 1,024, launching 600 sleep 1 processes with pipes at once produced 262 OSError: [Errno 24] Too many open files. And CPU-bound processes beyond the core count just time-slice: 192 CPU-bound jobs on a 24-core machine, all at once, took 3.99 s each (median) and the first result arrived after 1.10 s; with a limit of 24, each took 0.63 s, the first finished at 0.51 s, and half were done by 3.04 s instead of 4.27 s — at a total cost of 5.59 s against 4.55 s. This guide bounds subprocess concurrency and chooses the bound.

Prerequisites

1. Count what each subprocess costs

A subprocess with captured stdout and stderr holds descriptors in your process for the pipes, plus the child's own resources. With no bound, a batch multiplies all of it:

async def run(*cmd: str) -> bytes:
    proc = await asyncio.create_subprocess_exec(
        *cmd, stdout=asyncio.subprocess.PIPE, stderr=asyncio.subprocess.PIPE,
    )
    out, err = await proc.communicate()
    return out

# 600 at once under ulimit -n 1024: 262 failed with "Too many open files"
results = await asyncio.gather(*(run("sleep", "1") for _ in range(600)), return_exceptions=True)

The descriptor limit is per process and shared with everything else: sockets, database connections, log files. Hitting it does not only fail the extra subprocesses — any other code that opens a file or socket at that moment fails too, which is how a batch job takes down an unrelated health check. Memory is the other shared cost: each child has its own address space, so 200 Python children at 10–30 MiB each is gigabytes.

Verify: ls /proc/<pid>/fd | wc -l during a batch stays well below ulimit -n.

Launching 600 subprocesses with pipes, fd limit 1,024 2 horizontal bars comparing all 600 at once with the others. Launching 600 subprocesses with pipes, fd limit 1,024 all 600 at once 262 failed semaphore limit 100 0 failed, 6.1 s total Python 3.14 on Linux, sleep 1 with stdout and stderr pipes, RLIMIT_NOFILE soft limit 1,024. Unbounded spawning fails part of the batch and anything else that needs a descriptor.

2. Bound concurrency with a semaphore

Acquire a slot before starting each subprocess and release it when the process has exited and been reaped:

class SubprocessPool:
    def __init__(self, limit: int) -> None:
        self.slots = asyncio.Semaphore(limit)

    async def run(self, *cmd: str, timeout: float | None = None) -> tuple[int, bytes, bytes]:
        async with self.slots:                                   # wait for a free slot
            proc = await asyncio.create_subprocess_exec(
                *cmd, stdout=asyncio.subprocess.PIPE, stderr=asyncio.subprocess.PIPE,
            )
            try:
                async with asyncio.timeout(timeout):
                    out, err = await proc.communicate()
            except TimeoutError:
                proc.kill()
                await proc.communicate()                         # reap and drain the pipes
                raise
            return proc.returncode, out, err


pool = SubprocessPool(limit=os.cpu_count())
results = await asyncio.gather(*(pool.run("convert", f) for f in files), return_exceptions=True)

The slot is held until the process has been reaped — with communicate(), because a plain wait() after a kill can block on undrained pipes — so a slot never counts a process that is still running. Share one pool across the whole application, not one per request — otherwise the limit applies per request and the total is unbounded again. Kill and reap on timeout, and for commands that spawn children, kill the whole group as in killing subprocess trees on cancellation.

Verify: under load, the number of child processes never exceeds the limit (pgrep -P <pid> | wc -l).

3. Choose the limit from what the processes do

The right number depends on the resource each process mostly uses:

import os

CPU_BOUND = os.cpu_count()                      # compilers, encoders, image tools
IO_BOUND = 4 * os.cpu_count()                   # network tools, waiting on remote services
MEMORY_BOUND = int(available_memory_mib() * 0.7 // 400)    # e.g. ~400 MiB per headless browser
DISK_BOUND = 4                                  # tools that stream large files on one disk

Measured with 192 CPU-bound jobs on 24 cores: with no limit, every job ran for about 3.99 s because 192 processes time-sliced 24 cores, and the first results appeared only after 1.10 s. With a limit of 24, each job ran in 0.63 s, results started at 0.51 s and half the batch was done at 3.04 s. Total time was higher (5.59 s against 4.55 s) because the last wave ran with idle cores; that tail is the price of a bound. When results are consumed as they arrive — written to storage, returned to users — the steady flow is worth more than the slightly shorter total.

Verify: for CPU-bound tools, CPU utilization sits near 100% with the limit at the core count, and per-job run time stays close to the single-job time.

192 CPU-bound subprocesses on 24 cores A grid of 2 rows by 5 columns. 192 CPU-bound subprocesses on 24 cores limit run time per job first result half done all done none (192 at once) 3.99 s 1.10 s 4.27 s 4.55 s 24 (= cores) 0.63 s 0.51 s 3.04 s 5.59 s A bound trades a little total time for fast, steady results and bounded memory.

4. Stream work into the pool instead of gathering everything

gather over a large list creates every coroutine up front; with a semaphore inside, they mostly wait, but they all exist. For large or open-ended inputs, use a fixed number of workers pulling from a queue:

async def process_files(paths, limit: int) -> None:
    queue: asyncio.Queue[str] = asyncio.Queue(maxsize=limit * 2)

    async def worker() -> None:
        while True:
            path = await queue.get()
            try:
                code, out, err = await pool.run("thumbnail", path, timeout=60)
                if code != 0:
                    log.warning("thumbnail failed for %s: %s", path, err.decode()[-500:])
            finally:
                queue.task_done()

    async with asyncio.TaskGroup() as tg:
        workers = [tg.create_task(worker()) for _ in range(limit)]
        for path in paths:                    # can be a generator over millions of files
            await queue.put(path)
        await queue.join()
        for w in workers:
            w.cancel()

Memory now depends on the worker count, not on the number of inputs, and the producer waits when the queue is full. The pattern is the same one used for HTTP downloads in downloading many URLs concurrently with progress.

Verify: process a directory with 100,000 files; memory stays flat and the number of children never exceeds the limit.

5. Share the limit across the application

In a service, subprocesses are often started from request handlers. A global pool caps the total no matter how many requests arrive; requests beyond capacity wait or are refused:

@asynccontextmanager
async def lifespan(app):
    app.state.subprocs = SubprocessPool(limit=os.cpu_count())
    yield


async def render_pdf(request):
    pool: SubprocessPool = request.app.state.subprocs
    if pool.slots.locked():
        return JSONResponse({"error": "busy"}, status_code=503, headers={"Retry-After": "5"})
    code, out, err = await pool.run("wkhtmltopdf", "-", "-", timeout=30)
    return Response(out, media_type="application/pdf")

Refusing with 503 when all slots are busy keeps request latency honest; waiting is fine for batch callers but turns into piled-up requests for interactive ones. With multiple worker processes, each has its own pool, so the machine-wide total is workers × limit — divide accordingly, as for connection pools in why connection pools are per process.

Verify: a load test above capacity produces fast 503s rather than growing latency, and the machine's process count stays bounded.

What should limit these subprocesses? A decision on What does each subprocess mostly use with 4 outcomes. What should limit these subprocesses? What does each subprocess mostly use? CPU (encode, compile) limit = cores 0.63 s vs 3.99 s per job memory (browsers, ML) memory / per-process size avoid OOM network or remote waits a few x cores fds still bounded started from requests one shared pool 503 when full Every subprocess batch needs a bound; the resource decides the number.

Verification

Subprocess concurrency is bounded when:

  • All subprocesses go through one shared pool with an explicit limit.
  • The limit fits the dominant resource: cores, memory, descriptors or disk.
  • Large inputs are streamed to a fixed set of workers.
  • Timeouts kill and reap processes so slots are always returned.

Diagnostic Hook: export the pool's active count, waiters and per-job run time. Run time rising with active count means the limit is above what the machine can run in parallel; waiters growing while run time is flat means the machine has headroom and the limit can rise.

Pitfalls & edge cases

  • Unbounded gather over subprocesses. Measured: 262 of 600 failed on descriptors.
  • One pool per request. The limit no longer limits anything.
  • A limit far above cores for CPU-bound tools. Every job runs six times slower.
  • Forgetting to reap on timeout. Slots and zombie entries leak.

Frequently Asked Questions

How do I limit the number of concurrent subprocesses in asyncio?

Wrap each create_subprocess_exec and its communicate or wait in async with semaphore, using one asyncio.Semaphore shared across the application, sized to the resource the processes use.

Why do I get 'Too many open files' when starting many subprocesses?

Each subprocess with captured output uses several file descriptors in the parent. In testing, 600 at once under a limit of 1,024 descriptors failed 262 times; a limit of 100 concurrent processes fixed it.

How many CPU-bound subprocesses should run at once?

About one per core. In testing on 24 cores, 192 at once made each job take 3.99 s, while a limit of 24 kept each at 0.63 s and delivered results steadily.

Should a web service run subprocesses per request?

Only through a shared, bounded pool, refusing with 503 when it is full, so a traffic spike cannot fork an unbounded number of processes.