Skip to content

Choosing Process Start Methods

When an asyncio service offloads CPU work to a ProcessPoolExecutor, each worker process is created by one of three start methods, and the choice changes start-up time, memory, and whether the child can deadlock. Python 3.14 changed the default on Linux from fork to forkserver. Measured on Python 3.14.4 with a parent that had imported FastAPI, SQLAlchemy, Pydantic, httpx and aiohttp and held a 3-million-element list: a pool of 8 workers was ready in 39 ms with fork, 332 ms with forkserver and 345 ms with spawn. Creating further pools in the same process took 18–19 ms with fork, 8–11 ms with forkserver — whose one-off cost was starting its server — and 355–357 ms each time with spawn. The children's proportional memory (PSS) totalled 167 MiB with fork, 51–59 MiB with forkserver and 465 MiB with spawn, where every child had re-imported the parent's modules. Only fork children could see the parent's list. And forking while a second thread was running raised "DeprecationWarning: This process ... is multi-threaded, use of fork() may lead to deadlocks in the child." This guide explains which to use.

Prerequisites

1. Know what each method does

The three methods differ in how a worker process gets its starting state:

import multiprocessing as mp
from concurrent.futures import ProcessPoolExecutor

pool = ProcessPoolExecutor(8, mp_context=mp.get_context("fork"))        # copy of the parent
pool = ProcessPoolExecutor(8, mp_context=mp.get_context("forkserver"))  # forked from a clean server
pool = ProcessPoolExecutor(8, mp_context=mp.get_context("spawn"))       # fresh interpreter

fork copies the parent process as it is at that moment — every module, every object, every thread's held locks — using copy-on-write pages. spawn starts a fresh interpreter that imports the main module and unpickles the function it is asked to run. forkserver starts one clean server process once, and forks each worker from it; workers are copies of that server, not of your service. Measured: fork children could read the parent's 3-million-element list, and forkserver and spawn children could not. Anything a worker needs must therefore be passed as an argument, loaded in an initialiser, or imported at module level, as in initializing process pool workers with expensive state.

Verify: you know which start method your pools use — mp.get_start_method() reported forkserver on Python 3.14.4 on Linux — and your worker functions do not rely on parent state they cannot see.

2. Avoid fork in a process with threads

An asyncio service is rarely single-threaded: the default executor, database drivers, logging handlers and telemetry exporters run threads. fork copies only the thread that called it, along with every lock as it was — including locks held by the threads that were not copied, which no one in the child will ever release:

threading.Thread(target=background_work, daemon=True).start()
ctx = mp.get_context("fork")
ctx.Process(target=work, args=(10,)).start()
# DeprecationWarning: This process (pid=...) is multi-threaded,
# use of fork() may lead to deadlocks in the child.

Measured: with a second thread running, forking emitted exactly that DeprecationWarning. The deadlock itself is intermittent — it needs another thread to hold a lock, such as the logging module's or an allocator's, at the instant of the fork — which is why it passes tests and hangs in production. This is the reason Python 3.14 moved Linux to forkserver by default, and why fork should be chosen only for processes that are known to be single-threaded when the pool starts. Running tests with -W error::DeprecationWarning turns the warning into a failure that names the problem.

Verify: your service never forks after starting threads, or the warning is promoted to an error in tests.

Pool of 8 workers, parent with FastAPI, SQLAlchemy, Pydantic, httpx and aiohttp imported A grid of 3 rows by 6 columns. Pool of 8 workers, parent with FastAPI, SQLAlchemy, Pydantic, httpx and aiohttp imported method first pool later pools children PSS sees parent objects safe with threads fork 39 ms 18-19 ms 167 MiB yes no (DeprecationWarning) forkserver 332 ms 8-11 ms 51-59 MiB no yes spawn 345 ms 355-357 ms 465 MiB no yes Python 3.14.4 on Linux; PSS from /proc/<pid>/smaps_rollup, summed over 8 children.

3. Measure memory with PSS, not RSS

Memory per worker is easy to misread. Pages shared between a forked child and its parent count fully in each child's RSS, so summing RSS overstates the cost; PSS divides shared pages among their sharers:

def mem(pid: int) -> dict:
    out = {}
    for line in open(f"/proc/{pid}/smaps_rollup"):
        key, value, *_ = line.split()
        if key in ("Rss:", "Pss:", "Private_Clean:", "Private_Dirty:"):
            out[key.rstrip(":")] = int(value) / 1024          # MiB
    return out

Measured across 8 children: with fork, RSS summed to 1,423 MiB, but PSS was 167 MiB and private memory only 13 MiB — almost everything was shared with the parent. With spawn, RSS was 613 MiB, PSS 465 MiB and private 454 MiB: each child had imported the same libraries independently, about 57 MiB apiece. With forkserver, PSS was 51–59 MiB, since workers were forked from a small server and shared its pages. Copy-on-write sharing under fork also erodes over time — Python's reference counting writes to objects the child merely reads, so pages get copied — which is another reason the measured advantage of fork is smaller in long-running workers than at start-up.

Verify: worker memory is reported as PSS or private memory, never as summed RSS.

4. Pay forkserver's start-up once

forkserver's first pool is slow because it starts the server process, which imports the modules the workers will need. Later pools fork from the already-running server and are the fastest of the three:

ctx = mp.get_context("forkserver")
ctx.set_forkserver_preload(["myservice.cpu_tasks"])     # import heavy modules once, in the server

pool = ProcessPoolExecutor(8, mp_context=ctx)           # create at start-up, reuse for the process's life

Measured in one process creating three pools in turn: forkserver took 299.8 ms for the first and 7.7 and 9.3 ms for the next two; spawn took 350–357 ms every time; fork 17.7–19.0 ms every time. Explicitly preloading the heavy modules made no difference here — 306.3 ms, then 9.0 and 11.0 ms — because the main module and its imports were already being loaded by the server on Python 3.14. Create the pool once during application start-up, before taking traffic, so the one-off cost is not paid by a request; recreating pools per request is wasteful with any method and ruinous with spawn.

Verify: the pool is created once at start-up, and no request path creates a pool.

Time until 8 workers have run a task, three pools in a row 6 horizontal bars comparing forkserver, 1st pool with the others. Time until 8 workers have run a task, three pools in a row forkserver, 1st pool 299.8 ms forkserver, 2nd pool 7.7 ms fork, 1st pool 19.0 ms fork, 2nd pool 18.9 ms spawn, 1st pool 350.1 ms spawn, 2nd pool 355.0 ms Parent imports FastAPI, SQLAlchemy, Pydantic, httpx and aiohttp at module level. forkserver pays once; spawn pays every time.

5. Choose deliberately and say so in code

Set the start method explicitly where the pool is created, rather than relying on the platform default — it differs between Linux, macOS and Windows, and changed on Linux in Python 3.14:

import multiprocessing as mp

def make_cpu_pool(workers: int) -> ProcessPoolExecutor:
    # forkserver: safe with threads, small workers, fast after the first pool.
    ctx = mp.get_context("forkserver")
    return ProcessPoolExecutor(workers, mp_context=ctx, max_tasks_per_child=1000)

Choose forkserver for asyncio services on Linux: they have threads, and it is the only method that is both thread-safe and cheap after start-up. Choose spawn where forkserver is unavailable (Windows) or when workers must not inherit anything at all. Choose fork only for single-threaded batch scripts that rely on sharing large read-only data with workers, and test them with the deprecation warning as an error. Whichever you choose, the worker function and everything it needs must be importable from a module, not defined in __main__ under if __name__ == "__main__": — spawn and forkserver import it by name. Recycling long-lived workers is the subject of recycling process pool workers.

Verify: every pool states its start method explicitly, and the choice is documented next to it.

Which start method for this pool? A decision on What kind of program creates the pool with 4 outcomes. Which start method for this pool? What kind of program creates the pool? async service on Linux (threads) forkserver 8-11 ms after first pool Windows, or nothing may be inherited spawn ~350 ms per pool single-threaded batch, big shared data fork, warning as error sees parent objects any set mp_context explicitly defaults differ by platform Python 3.14 made forkserver the Linux default for these reasons.

Verification

Process start methods are chosen well when:

  • Every pool sets its mp_context explicitly.
  • Nothing forks after threads have started, or the warning fails tests.
  • Worker memory is measured as PSS, and the pool is created once at start-up.
  • Worker functions are importable, with state passed in or loaded by an initialiser.

Diagnostic Hook: if a process pool's workers occasionally hang at start-up or on their first task, with no exception and no CPU use, check the start method. A child stuck on a lock inherited through fork from a thread that no longer exists is the classic signature, and switching to forkserver removes it.

Pitfalls & edge cases

  • fork in a threaded service. Measured: the multi-threaded fork DeprecationWarning.
  • Summing RSS across workers. Measured: 1,423 MiB summed against 167 MiB PSS.
  • spawn pools created repeatedly. Measured: about 350 ms each time.
  • Relying on parent globals in workers. Only fork children saw them.

Frequently Asked Questions

What is the default multiprocessing start method in Python 3.14?

On Linux it is now forkserver; mp.get_start_method() reported forkserver on Python 3.14.4. macOS and Windows default to spawn.

Is fork safe to use in an asyncio application?

Not once threads are running: forking with a second thread active raised "This process ... is multi-threaded, use of fork() may lead to deadlocks in the child." Use forkserver.

Which start method starts process pools fastest?

In testing, fork took 18 to 39 ms per pool; forkserver took 300 ms for its first pool and 8 to 11 ms afterwards; spawn about 350 ms every time.

Which start method uses the least memory?

forkserver: children totalled 51 to 59 MiB PSS, against 167 MiB for fork and 465 MiB for spawn, where each child re-imported the parent's libraries.