Skip to content

Recycling Process Pool Workers

Worker processes in a long-lived pool accumulate whatever their tasks leave behind: caches that grow, C extensions that leak, fragmented heaps that never shrink. Recycling — replacing each worker after a number of tasks — caps that growth. ProcessPoolExecutor gained a max_tasks_per_child parameter for this in Python 3.11, and in testing it was the option to be careful with. Measured with a task that retained 2 MiB per call, 400 calls on 4 workers: without recycling, the workers' combined PSS peaked at 811 MiB. With max_tasks_per_child=1, it stayed at 22 MiB, but throughput fell from 1,753 to 648 tasks per second — and for a task that leaked nothing, from 6,027 to 783. With max_tasks_per_child set to 2 or more, the executor hung on Python 3.12.13, 3.13.14, 3.14.4 and 3.14.6 as soon as a worker reached its limit: workers exited, no replacements started, and the waiting futures never completed. multiprocessing.Pool(maxtasksperchild=25) bridged into asyncio kept PSS at 153 MiB at 2,029 tasks/s (5,544 for clean tasks), and replacing the whole executor every 100 tasks kept it at 163 MiB at 1,819 tasks/s. This guide shows how to recycle safely.

Prerequisites

1. Measure whether workers grow at all

Recycling has a cost, so first confirm there is growth to cap. Sample the workers' proportional set size while the pool runs real work:

def workers_pss(pool) -> float:
    total = 0.0
    for pid in list(pool._processes or {}):
        with open(f"/proc/{pid}/smaps_rollup") as f:
            for line in f:
                if line.startswith("Pss:"):
                    total += int(line.split()[1]) / 1024
    return total                                          # MiB

Measured with a task that appended a 2 MiB buffer to a module-level list on each call: after 400 calls on 4 workers, PSS peaked at 811 MiB — 400 × 2 MiB, split across the workers — and kept rising with every task. In a real service the growth is usually slower and less obvious: a cache keyed by input, a library holding references, an allocator that does not return memory to the system. A steady upward slope in worker memory under a stable workload is the signal; a plateau means recycling is not needed.

Verify: you have a graph of worker memory over hours of production traffic, and it either plateaus or you know its growth rate.

2. Know that max_tasks_per_child above 1 can hang

The obvious fix is the executor's own parameter:

pool = ProcessPoolExecutor(4, max_tasks_per_child=10)

Measured: with any value of 2 or more and more tasks submitted than workers × max_tasks_per_child, the executor stopped making progress once the first worker reached its limit — 41 tasks with a limit of 10 and 4 workers was enough — on Python 3.12.13, 3.13.14, 3.14.4 and 3.14.6, with both forkserver and spawn. The main thread waited on a future, the executor's manager thread waited for results, and the process list showed the resource tracker and forkserver but no workers. The cause is visible in concurrent/futures/process.py: each completed task releases an "idle worker" semaphore, and when a worker exits at its limit, _adjust_process_count() first takes one of those stale permits and returns without spawning a replacement. With max_tasks_per_child=1, no permits accumulate, and replacements are always spawned — which is why only that value worked. Test this on your exact interpreter before relying on it; a fix may land in a later release.

Verify: a test that submits more than workers × max_tasks_per_child tasks completes on your Python version.

ProcessPoolExecutor with max_tasks_per_child, forkserver A grid of 5 rows by 4 columns. ProcessPoolExecutor with max_tasks_per_child, forkserver max_tasks_per_child workers tasks result 1 2 or 4 20 completed 10 4 40 (no replacement needed) completed 10 4 41 hung 2 / 3 / 5 2-4 20-30 hung 10 4 400 (spawn) hung Reproduced on Python 3.12.13, 3.13.14, 3.14.4 and 3.14.6.

3. Use multiprocessing.Pool's maxtasksperchild from asyncio

multiprocessing.Pool has had maxtasksperchild for many years and replaced workers correctly in these tests. It has no asyncio integration, but its callbacks can resolve futures on the loop:

ctx = mp.get_context("forkserver")
pool = ctx.Pool(4, maxtasksperchild=25)

def submit(loop, fn, *args) -> asyncio.Future:
    fut = loop.create_future()
    pool.apply_async(
        fn, args,
        callback=lambda r: loop.call_soon_threadsafe(fut.set_result, r),
        error_callback=lambda e: loop.call_soon_threadsafe(fut.set_exception, e),
    )
    return fut

results = await asyncio.gather(*(submit(loop, leaky, i) for i in range(400)))

The callbacks run on the pool's result-handler thread, so they hand results to the loop with call_soon_threadsafe. Measured with the leaking task: 2,029 tasks per second and a workers' PSS peak of 153 MiB, against 811 MiB without recycling; with a task that leaked nothing, 5,544 tasks per second, close to the 6,027 of an unrecycled ProcessPoolExecutor. A cancelled asyncio future does not stop the task in the pool — the result arrives and is discarded — so for cancellable work prefer the executor-based approach in step 4, or the techniques in cancelling long-running work in a process pool.

Verify: under sustained load, worker memory stays bounded and tasks complete past many recycling cycles.

Workers' PSS peak, 400 tasks retaining 2 MiB each 4 horizontal bars comparing no recycling with the others. Workers' PSS peak, 400 tasks retaining 2 MiB each no recycling 811 MiB new executor every 100 tasks 163 MiB mp.Pool, maxtasksperchild=25 153 MiB max_tasks_per_child=1 22 MiB Throughput with a leak-free task: 6,027 / 4,541 / 5,544 / 783 tasks/s respectively. Recycling every task is the smallest and by far the slowest.

4. Or rotate the whole executor

Another way to recycle, entirely within ProcessPoolExecutor, is to replace the executor itself after a number of tasks: new work goes to a fresh pool, and the old one shuts down once its tasks finish:

class RotatingPool:
    def __init__(self, workers: int, tasks_per_pool: int):
        self.workers, self.tasks_per_pool = workers, tasks_per_pool
        self.pool, self.used = self._new(), 0

    def _new(self):
        return ProcessPoolExecutor(self.workers, mp_context=mp.get_context("forkserver"))

    async def run(self, fn, *args):
        loop = asyncio.get_running_loop()
        if self.used >= self.tasks_per_pool:
            old, self.pool, self.used = self.pool, self._new(), 0
            loop.run_in_executor(None, old.shutdown)            # waits for old tasks, off the loop
        self.used += 1
        return await loop.run_in_executor(self.pool, fn, *args)

Measured with a new executor every 100 tasks: 1,819 tasks per second and a PSS peak of 163 MiB with the leaking task; 4,541 tasks per second with a clean one. Rotation costs a pool start-up each cycle — cheap with forkserver after its first pool, as measured in choosing process start methods — and briefly runs two sets of workers while the old pool drains, so size tasks_per_pool large enough that this overlap is rare.

Verify: old executors shut down after rotation, and the number of worker processes stays at most twice the pool size.

5. Choose the recycling interval from the growth rate

Recycling every task gives the smallest workers and pays a process start per task: with max_tasks_per_child=1, a leak-free task ran at 783 tasks per second instead of 6,027. Choose the interval so workers are replaced well before their growth matters:

GROWTH_PER_TASK_MIB = 0.05        # measured in production
BUDGET_PER_WORKER_MIB = 200
TASKS_PER_CHILD = int(BUDGET_PER_WORKER_MIB / GROWTH_PER_TASK_MIB / 2)   # recycle at half budget

Add jitter to the interval so workers do not all restart at the same moment, and log each recycle with the worker's final memory, so a later change in the growth rate is visible. If growth is caused by a cache inside the worker, bounding the cache — as in bounding in-memory async caches by size — is a better fix than recycling around it.

Verify: the recycling interval keeps worker memory under its budget with a measured throughput cost.

How should these workers be recycled? A decision on Does worker memory grow with 4 outcomes. How should these workers be recycled? Does worker memory grow? no, it plateaus no recycling 6,027 tasks/s yes mp.Pool(maxtasksperchild) or rotate the executor 153-163 MiB peak must use ProcessPoolExecutor's own option max_tasks_per_child=1 783 tasks/s max_tasks_per_child >= 2 test first hung on 3.12-3.14 Fix the leak if you can; recycle around it if you cannot.

Verification

Worker recycling is safe and worth its cost when:

  • Worker memory growth is measured, and recycling is used only where it grows.
  • The recycling mechanism is tested past several cycles on the exact Python version in use.
  • The interval is derived from the growth rate, with jitter, and its throughput cost is known.
  • Each recycle is logged with the worker's final memory.

Diagnostic Hook: if a process pool stops completing tasks while CPU use drops to zero and the worker processes have disappeared, check for max_tasks_per_child. A pool that hangs exactly after its first workers reach their task limit is the executor failing to spawn replacements, not a slow task.

Pitfalls & edge cases

  • max_tasks_per_child of 2 or more. Measured: hung on Python 3.12 through 3.14.6.
  • max_tasks_per_child=1 for throughput-sensitive work. Measured: 783 instead of 6,027 tasks/s.
  • Cancelling asyncio futures backed by multiprocessing.Pool. The task still runs.
  • Recycling to hide a cache. Bound the cache instead.

Frequently Asked Questions

How do I restart ProcessPoolExecutor workers after N tasks?

Python 3.11+ has max_tasks_per_child, but in testing any value of 2 or more hung once a worker had to be replaced, on Python 3.12.13 to 3.14.6. Use max_tasks_per_child=1, multiprocessing.Pool(maxtasksperchild=N), or replace the executor periodically.

Why does my ProcessPoolExecutor hang with max_tasks_per_child?

When a worker exits at its limit, the executor takes a stale idle-worker permit instead of spawning a replacement; once all workers have exited, nothing runs the queued tasks.

What does worker recycling cost?

With max_tasks_per_child=1, a short task ran at 783 tasks/s instead of 6,027; with multiprocessing.Pool recycling every 25 tasks, 5,544.

Do I need to recycle process pool workers?

Only if their memory grows under a stable workload. A task retaining 2 MiB per call grew 4 workers to 811 MiB in 400 calls; without growth, recycling only costs throughput.