Skip to content

Initializing Process Pool Workers with Expensive State

Many CPU-bound tasks need something expensive before they can do anything: a machine-learning model loaded from disk, a compiled regex set, a spatial index, a tokenizer. Loading it inside the task function means paying that cost on every call. A ProcessPoolExecutor initializer runs a function once in each worker process as it starts, so the state is built once per worker and reused by every task it runs. Measured with a stand-in load of 300 ms and eight workers, 200 calls that loaded the state per call took 7.6 s; with the load moved to an initializer, the same 200 calls — pool start-up included — took 0.32 s. This guide sets up initializers from asyncio, explains why they matter even more on Python 3.14, and covers refreshing the state.

Prerequisites

1. Load state in an initializer, read it from a global

The initializer and the task function run in the same worker process, so a module-level global is the natural place to keep the state:

import asyncio
import time
from concurrent.futures import ProcessPoolExecutor

_model = None                         # one per worker process


def load_model(path: str):
    time.sleep(0.3)                   # stand-in for reading weights, building an index, ...
    return {"w": 3, "path": path}


def init_worker(path: str) -> None:
    global _model
    _model = load_model(path)


def predict(x: float) -> float:
    return _model["w"] * x            # no loading here


async def main() -> None:
    loop = asyncio.get_running_loop()
    with ProcessPoolExecutor(max_workers=8, initializer=init_worker,
                             initargs=("/models/v3.bin",)) as pool:
        results = await asyncio.gather(*(loop.run_in_executor(pool, predict, i) for i in range(200)))


if __name__ == "__main__":
    asyncio.run(main())

Measured: 0.32 s for 200 calls including starting eight workers that each loaded once, against 7.6 s when predict loaded the model itself. Pass configuration through initargs rather than relying on globals in the parent: with the forkserver and spawn start methods, the parent's globals are not inherited.

Verify: count loads with a counter in the initializer — one per worker, regardless of how many tasks run.

200 calls that need a 300 ms setup 2 horizontal bars comparing load inside every call with the others. 200 calls that need a 300 ms setup load inside every call 7.6 s load once per worker 0.32 s 8 worker processes, Python 3.14, setup simulated with a 300 ms sleep; times include pool start-up. The initializer turns a per-call cost into a per-worker cost.

2. Understand what the start method does to state

The three start methods treat parent state differently, and the default changed on Linux in Python 3.14:

Start method Parent globals in workers Default on Linux
fork copied at fork time until 3.13
forkserver not inherited; modules re-imported 3.14+
spawn not inherited; fresh interpreter macOS and Windows

Code that loaded a model at import time in the parent and relied on fork to share it with workers silently changes behaviour on 3.14: each worker re-imports the module and loads the model again, or finds None if the load was guarded by if __name__ == "__main__". An explicit initializer behaves the same under every start method, which is the main reason to use one even where fork-copying used to "just work".

Fork's copy-on-write sharing was also less free than it looks: CPython's reference counting writes to object headers, so pages holding Python objects get copied as soon as workers touch them. Large NumPy arrays benefit from copy-on-write; large dicts of Python objects mostly do not.

Verify: run your pool with mp_context=multiprocessing.get_context("spawn") in a test; if it breaks, it depended on fork inheritance.

3. Handle initializer failures

If the initializer raises, the worker dies — and ProcessPoolExecutor marks the whole pool broken. Every pending and future call fails with BrokenProcessPool:

from concurrent.futures.process import BrokenProcessPool


def init_worker(path: str) -> None:
    global _model
    try:
        _model = load_model(path)
    except Exception:
        log.exception("worker init failed for %s", path)
        raise                                  # the pool becomes unusable


async def predict_safe(pool, x):
    loop = asyncio.get_running_loop()
    try:
        return await loop.run_in_executor(pool, predict, x)
    except BrokenProcessPool:
        raise ServiceUnavailable("model workers failed to start") from None

Fail early instead: load the state once in the parent at startup — or at least validate the path and file — so a missing model aborts the deploy rather than breaking the pool on first use. Recovering a broken pool by recreating it is covered in recovering from BrokenProcessPool in async services.

Verify: point the initializer at a missing file; startup fails with a clear error, or the first call raises BrokenProcessPool that your service maps to a 503.

What happens when a worker starts A sequence of 6 messages between 3 participants. What happens when a worker starts event loop pool worker process run_in_executor(predict, 1) start worker initializer: load model (300 ms) predict(1) using _model run_in_executor(predict, 2) predict(2): no load Only the first task in each worker waits for the load; every later one runs immediately.

4. Refresh state without restarting the service

Long-lived state goes stale: a new model version, an updated index. Options, from simplest to most precise:

async def swap_pool(app, new_path: str) -> None:
    """Start a new pool with the new state, switch traffic, then retire the old one."""
    loop = asyncio.get_running_loop()
    new_pool = ProcessPoolExecutor(max_workers=8, initializer=init_worker, initargs=(new_path,))
    await asyncio.gather(*(loop.run_in_executor(new_pool, predict, 0) for _ in range(8)))  # warm
    old_pool, app.pool = app.pool, new_pool
    await loop.run_in_executor(None, old_pool.shutdown, True)                           # drain

Building a second pool, warming it, swapping the reference, and shutting down the old one gives a clean cut-over with no request seeing a half-loaded model, at the cost of briefly running two pools' worth of memory. The alternative — max_tasks_per_child, which recycles each worker after a number of tasks so it re-runs the initializer — sounds simpler, but test it carefully in your environment: on the Linux host used for this guide (kernel 7.0, CPython 3.12 to 3.14), pools with max_tasks_per_child hung as soon as a worker needed replacing, even in a two-worker reproduction. The explicit swap avoids that machinery entirely.

Verify: during a swap under load, no call fails and every response after the swap comes from the new state (include the version in the result).

5. Size the pool around the state's memory

Each worker holds its own copy of the state, so memory is workers × state size plus the parent:

import os


def workers_for(state_bytes: int, memory_limit: int, reserve: float = 0.3) -> int:
    budget = memory_limit * (1 - reserve)                   # leave room for the parent and spikes
    by_memory = max(1, int(budget // state_bytes))
    by_cpu = os.process_cpu_count() or 1
    return min(by_memory, by_cpu)

A 1.5 GB model on a 4 GB container fits two workers, whatever the CPU count. When memory, not CPU, limits the worker count, consider sharing read-only arrays through shared memory, as in sharing large arrays with shared memory, or running the model in a separate inference service that the async service calls.

Verify: peak RSS across the parent and workers stays under the container limit with headroom, measured during a pool swap when two pools briefly coexist.

Where should expensive per-task state live? A decision on How big is the state per worker with 3 outcomes. Where should expensive per-task state live? How big is the state per worker? fits several times over initializer + global once per worker large read-only arrays shared memory one copy for all does not fit N times separate inference service call it over the network Load once per worker by default; share or move it out only when memory forces you to.

Verification

Worker initialization is right when:

  • State loads once per worker, verified by a counter, never per call.
  • Behaviour is identical under fork, forkserver and spawn.
  • Initializer failures surface at startup or as a clear 503, never as silent hangs.
  • Refreshing state uses a warmed replacement pool, and memory has room for both during the swap.

Diagnostic Hook: log each initializer run with its duration, the worker pid and the state version, and export state version per worker as a gauge. During a model rollout, the gauge shows exactly when every worker serves the new version; an initializer duration that creeps up across releases is cold-start latency your users will feel after every deploy.

Pitfalls & edge cases

  • Loading state at import time. Re-runs in every worker under forkserver and spawn, and in the parent too.
  • Relying on fork inheritance. Python 3.14 changed the Linux default to forkserver; pass state through initargs.
  • Mutable shared state. Workers' globals are independent copies; writes in one worker are invisible to others.
  • Unbounded worker counts. Each worker multiplies the state's memory.

Frequently Asked Questions

How do I load a model once per process pool worker?

Pass initializer=func and initargs to ProcessPoolExecutor. The initializer runs once in each worker as it starts and stores the model in a module-level global that task functions read. In testing this took 200 calls from 7.6 s to 0.32 s.

Why does my process pool worker not see globals from the parent?

With the forkserver or spawn start methods, workers do not inherit the parent's memory; they re-import modules. Python 3.14 made forkserver the default on Linux. Pass state through the initializer instead.

What happens if a ProcessPoolExecutor initializer raises?

The worker exits and the pool is marked broken, so pending and later calls raise BrokenProcessPool. Validate the inputs at startup so failures happen before the pool is used.

How do I reload worker state without restarting the service?

Create a new pool whose initializer loads the new state, warm it, switch the reference used by callers, and shut the old pool down after it drains.