Shutting Down Multi-Process asyncio Servers¶
A multi-process server is a supervisor and its workers, and shutting it down means the signal reaching the supervisor, the supervisor stopping the workers, and each worker draining its own requests. The happy path works; the failure paths leave processes behind. Measured on Python 3.14 with uvicorn 0.54 running 4 workers and 40 requests of 2 seconds in flight: SIGTERM to the master let all 40 complete with 200, served by all 4 workers, and left 0 of the master's 5 child processes alive. With --timeout-graceful-shutdown 1, all 40 were cut off with 500 at the 1-second limit. SIGKILL to the master — what an out-of-memory kill or a crashed supervisor looks like — left all 5 children running: the 40 in-flight requests still completed, a new request 2 seconds later was served, and starting a fresh server on the same port failed with [Errno 98] address already in use. This guide makes both paths end with no stragglers.
Prerequisites¶
- uvicorn with
--workers, or another pre-fork server, on Linux. - Single-process shutdown, from enforcing a hard shutdown deadline.
- The topic overview, Graceful Shutdown & Signals.
1. Measure the supervised path¶
Start the server with several workers, put requests in flight, and send SIGTERM to the master only — as docker stop and Kubernetes do, to PID 1:
p = subprocess.Popen([sys.executable, "-m", "uvicorn", "app:app",
"--port", "8000", "--workers", "4"])
workers = children(p.pid) # from /proc/<pid>/task/<pid>/children
# ... 40 requests of 2 s each in flight ...
p.send_signal(signal.SIGTERM)
Measured: all 40 requests completed with 200, spread across all 4 workers; the master exited with code 0; and 3 seconds later none of its 5 children — the 4 workers and one helper process — were alive. The master forwarded the signal to each worker, each worker stopped accepting and drained its in-flight requests, and the master waited for them. This is the path to preserve.
Verify: a SIGTERM to the master with requests in flight completes every request and leaves no child processes.
2. Size the drain timeout to the slowest request¶
uvicorn's --timeout-graceful-shutdown bounds how long each worker drains before cancelling what is left. Measured with 2-second requests and a 1-second limit: all 40 requests were cancelled and the clients received 500 responses. Without the option, the workers waited as long as the requests needed. Neither extreme is right: an unbounded drain can exceed the platform's grace period and be killed anyway, and a short one fails requests that would have finished:
uvicorn app:app --workers 4 --timeout-graceful-shutdown 25 # p99.9 request time + margin, < platform grace
Set it from measured request durations — the slowest requests you are willing to let finish — and keep it below the platform's termination grace period, leaving time for the master to reap workers and exit. Long-lived connections such as WebSockets and streaming responses never "finish", so they need their own close logic, as in closing WebSocket connections on shutdown.
Verify: the drain timeout exceeds the p99.9 request duration and is below the platform's grace period.
3. See what happens when the master dies¶
The supervisor can die without running its shutdown code: an out-of-memory kill, a kill -9, a crash. Measured with SIGKILL to the master: the in-flight requests still completed — the workers never noticed — and all 5 children were still alive 3 seconds later, reparented to init. They kept their listening sockets: a new request 2 seconds after the kill was served with 200. And a replacement server started on the same port exited with code 3 and ERROR: [Errno 98] error while attempting to bind on address ('127.0.0.1', 58831): address already in use. A process supervisor that restarts the service in place will fail to start it until the orphans are found and killed, while the orphans run old code with no supervision.
def orphaned_workers(master_pid: int, worker_pids: list[int]) -> list[int]:
return [w for w in worker_pids if alive(w) and not alive(master_pid)]
Verify: after SIGKILL to the master in a test, count surviving children; any survivor means workers need a way to notice the master's death.
4. Make workers notice a dead parent¶
Two mechanisms close the gap. Run the server as a container's main process: when PID 1 of a container's PID namespace exits, the kernel kills every other process in it, which is the main reason this failure is rarer in containers than on bare hosts and under process supervisors. Inside the worker, ask the kernel for a signal when the parent dies:
import ctypes, signal
PR_SET_PDEATHSIG = 1
def die_with_parent(sig=signal.SIGTERM):
libc = ctypes.CDLL("libc.so.6", use_errno=True)
if libc.prctl(PR_SET_PDEATHSIG, sig) != 0:
raise OSError(ctypes.get_errno(), "prctl failed")
# call at the start of each worker process, before serving
Measured with a minimal parent and child: after SIGKILL to the parent, a child without the setting was still alive a second later; a child that had called prctl(PR_SET_PDEATHSIG, SIGTERM) received SIGTERM, ran its handler and exited. The setting is Linux-specific and is not inherited by a forked child, so call it in the worker itself, after the fork. A portable alternative is a worker task that polls os.getppid() every second and initiates its own graceful shutdown when the parent PID changes to 1 or to a subreaper. Either way, the orphaned worker then runs its normal drain, rather than serving indefinitely.
Verify: after SIGKILL to the master, every worker exits within the detection interval, and the port is free for a restart.
5. Shut down process pools inside workers too¶
Workers often own child processes of their own — a ProcessPoolExecutor for CPU work. Those are a second level of the same problem: shut them down in the worker's lifespan, or they outlive it. With Python 3.14's default forkserver start method, pool processes are not children of the worker, as measured in validating large request bodies off the loop, where six processes outlived a SIGTERM until the pool was shut down in the lifespan:
@asynccontextmanager
async def lifespan(app):
app.state.pool = ProcessPoolExecutor(max_workers=2)
yield
app.state.pool.shutdown(wait=True, cancel_futures=True)
After every change to the process structure, test shutdown the way it will happen in production: SIGTERM to PID 1 with requests in flight, then SIGKILL to the master, counting survivors each time.
Verify: after both tests, no process from the service remains, checked by PID rather than by name.
Verification¶
A multi-process server shuts down cleanly when:
- SIGTERM to the master drains every in-flight request and leaves no child processes.
- The drain timeout covers slow requests and fits inside the platform's grace period.
- Workers notice a dead master, through
PR_SET_PDEATHSIGor parent polling, or the platform kills the whole process group. - Nested process pools are shut down in each worker's lifespan.
Diagnostic Hook: when a service restart fails with "address already in use" after a crash, look for orphaned workers from the previous master — processes whose parent is now PID 1 and which still hold the port. A SIGKILL to the uvicorn master left all 5 children running and serving requests in this test.
Pitfalls & edge cases¶
- A drain timeout shorter than requests. Measured: 40 of 40 cut off with 500.
- Assuming workers die with the master. Measured: 5 of 5 survived SIGKILL.
- Restarting in place after a crash. Measured: bind failed, address in use.
- Forgetting pools inside workers. They outlive the worker as well.
Frequently Asked Questions¶
Does uvicorn --workers shut down gracefully on SIGTERM?
Yes: SIGTERM to the master let all 40 in-flight requests finish across 4 workers and left no child processes behind.
What happens to uvicorn workers if the master is killed?
They keep running. After SIGKILL to the master, all 5 children survived, still served requests, and blocked a restart on the same port.
How do I make worker processes exit when their parent dies?
On Linux, call prctl(PR_SET_PDEATHSIG, SIGTERM) in each worker, or poll os.getppid() and start a graceful shutdown when it changes.
What should --timeout-graceful-shutdown be set to?
Above your slowest normal request and below the platform grace period. At 1 s with 2 s requests, all 40 in-flight requests returned 500.
Related¶
- Graceful Shutdown & Signals — up to the topic overview.
- Stopping Kafka consumers gracefully — draining a different kind of worker.
- Resilience, Cancellation & Error Handling — the section overview.