Initializing Process Pool Workers with Expensive State¶
Many CPU-bound tasks need something expensive before they can do anything: a machine-learning model loaded from disk, a compiled regex set, a spatial index, a tokenizer. Loading it inside the task function means paying that cost on every call. A ProcessPoolExecutor initializer runs a function once in each worker process as it starts, so the state is built once per worker and reused by every task it runs. Measured with a stand-in load of 300 ms and eight workers, 200 calls that loaded the state per call took 7.6 s; with the load moved to an initializer, the same 200 calls — pool start-up included — took 0.32 s. This guide sets up initializers from asyncio, explains why they matter even more on Python 3.14, and covers refreshing the state.
Prerequisites¶
- Python 3.11+, stdlib only.
- Process pools from asyncio, from offloading CPU work with run_in_executor.
- Start methods — on Python 3.14 the default on Linux is forkserver, as noted in choosing chunksize for ProcessPoolExecutor work.
1. Load state in an initializer, read it from a global¶
The initializer and the task function run in the same worker process, so a module-level global is the natural place to keep the state:
import asyncio
import time
from concurrent.futures import ProcessPoolExecutor
_model = None # one per worker process
def load_model(path: str):
time.sleep(0.3) # stand-in for reading weights, building an index, ...
return {"w": 3, "path": path}
def init_worker(path: str) -> None:
global _model
_model = load_model(path)
def predict(x: float) -> float:
return _model["w"] * x # no loading here
async def main() -> None:
loop = asyncio.get_running_loop()
with ProcessPoolExecutor(max_workers=8, initializer=init_worker,
initargs=("/models/v3.bin",)) as pool:
results = await asyncio.gather(*(loop.run_in_executor(pool, predict, i) for i in range(200)))
if __name__ == "__main__":
asyncio.run(main())
Measured: 0.32 s for 200 calls including starting eight workers that each loaded once, against 7.6 s when predict loaded the model itself. Pass configuration through initargs rather than relying on globals in the parent: with the forkserver and spawn start methods, the parent's globals are not inherited.
Verify: count loads with a counter in the initializer — one per worker, regardless of how many tasks run.
2. Understand what the start method does to state¶
The three start methods treat parent state differently, and the default changed on Linux in Python 3.14:
| Start method | Parent globals in workers | Default on Linux |
|---|---|---|
| fork | copied at fork time | until 3.13 |
| forkserver | not inherited; modules re-imported | 3.14+ |
| spawn | not inherited; fresh interpreter | macOS and Windows |
Code that loaded a model at import time in the parent and relied on fork to share it with workers silently changes behaviour on 3.14: each worker re-imports the module and loads the model again, or finds None if the load was guarded by if __name__ == "__main__". An explicit initializer behaves the same under every start method, which is the main reason to use one even where fork-copying used to "just work".
Fork's copy-on-write sharing was also less free than it looks: CPython's reference counting writes to object headers, so pages holding Python objects get copied as soon as workers touch them. Large NumPy arrays benefit from copy-on-write; large dicts of Python objects mostly do not.
Verify: run your pool with mp_context=multiprocessing.get_context("spawn") in a test; if it breaks, it depended on fork inheritance.
3. Handle initializer failures¶
If the initializer raises, the worker dies — and ProcessPoolExecutor marks the whole pool broken. Every pending and future call fails with BrokenProcessPool:
from concurrent.futures.process import BrokenProcessPool
def init_worker(path: str) -> None:
global _model
try:
_model = load_model(path)
except Exception:
log.exception("worker init failed for %s", path)
raise # the pool becomes unusable
async def predict_safe(pool, x):
loop = asyncio.get_running_loop()
try:
return await loop.run_in_executor(pool, predict, x)
except BrokenProcessPool:
raise ServiceUnavailable("model workers failed to start") from None
Fail early instead: load the state once in the parent at startup — or at least validate the path and file — so a missing model aborts the deploy rather than breaking the pool on first use. Recovering a broken pool by recreating it is covered in recovering from BrokenProcessPool in async services.
Verify: point the initializer at a missing file; startup fails with a clear error, or the first call raises BrokenProcessPool that your service maps to a 503.
4. Refresh state without restarting the service¶
Long-lived state goes stale: a new model version, an updated index. Options, from simplest to most precise:
async def swap_pool(app, new_path: str) -> None:
"""Start a new pool with the new state, switch traffic, then retire the old one."""
loop = asyncio.get_running_loop()
new_pool = ProcessPoolExecutor(max_workers=8, initializer=init_worker, initargs=(new_path,))
await asyncio.gather(*(loop.run_in_executor(new_pool, predict, 0) for _ in range(8))) # warm
old_pool, app.pool = app.pool, new_pool
await loop.run_in_executor(None, old_pool.shutdown, True) # drain
Building a second pool, warming it, swapping the reference, and shutting down the old one gives a clean cut-over with no request seeing a half-loaded model, at the cost of briefly running two pools' worth of memory. The alternative — max_tasks_per_child, which recycles each worker after a number of tasks so it re-runs the initializer — sounds simpler, but test it carefully in your environment: on the Linux host used for this guide (kernel 7.0, CPython 3.12 to 3.14), pools with max_tasks_per_child hung as soon as a worker needed replacing, even in a two-worker reproduction. The explicit swap avoids that machinery entirely.
Verify: during a swap under load, no call fails and every response after the swap comes from the new state (include the version in the result).
5. Size the pool around the state's memory¶
Each worker holds its own copy of the state, so memory is workers × state size plus the parent:
import os
def workers_for(state_bytes: int, memory_limit: int, reserve: float = 0.3) -> int:
budget = memory_limit * (1 - reserve) # leave room for the parent and spikes
by_memory = max(1, int(budget // state_bytes))
by_cpu = os.process_cpu_count() or 1
return min(by_memory, by_cpu)
A 1.5 GB model on a 4 GB container fits two workers, whatever the CPU count. When memory, not CPU, limits the worker count, consider sharing read-only arrays through shared memory, as in sharing large arrays with shared memory, or running the model in a separate inference service that the async service calls.
Verify: peak RSS across the parent and workers stays under the container limit with headroom, measured during a pool swap when two pools briefly coexist.
Verification¶
Worker initialization is right when:
- State loads once per worker, verified by a counter, never per call.
- Behaviour is identical under fork, forkserver and spawn.
- Initializer failures surface at startup or as a clear 503, never as silent hangs.
- Refreshing state uses a warmed replacement pool, and memory has room for both during the swap.
Diagnostic Hook: log each initializer run with its duration, the worker pid and the state version, and export state version per worker as a gauge. During a model rollout, the gauge shows exactly when every worker serves the new version; an initializer duration that creeps up across releases is cold-start latency your users will feel after every deploy.
Pitfalls & edge cases¶
- Loading state at import time. Re-runs in every worker under forkserver and spawn, and in the parent too.
- Relying on fork inheritance. Python 3.14 changed the Linux default to forkserver; pass state through
initargs. - Mutable shared state. Workers' globals are independent copies; writes in one worker are invisible to others.
- Unbounded worker counts. Each worker multiplies the state's memory.
Frequently Asked Questions¶
How do I load a model once per process pool worker?
Pass initializer=func and initargs to ProcessPoolExecutor. The initializer runs once in each worker as it starts and stores the model in a module-level global that task functions read. In testing this took 200 calls from 7.6 s to 0.32 s.
Why does my process pool worker not see globals from the parent?
With the forkserver or spawn start methods, workers do not inherit the parent's memory; they re-import modules. Python 3.14 made forkserver the default on Linux. Pass state through the initializer instead.
What happens if a ProcessPoolExecutor initializer raises?
The worker exits and the pool is marked broken, so pending and later calls raise BrokenProcessPool. Validate the inputs at startup so failures happen before the pool is used.
How do I reload worker state without restarting the service?
Create a new pool whose initializer loads the new state, warm it, switch the reference used by callers, and shut the old pool down after it drains.
Related¶
- CPU-Bound Task Offloading — up to the topic overview.
- Cancelling long-running work in a process pool — stopping work these workers are doing.
- Concurrent Execution & Worker Patterns — the section overview.