Enforcing a Hard Shutdown Deadline¶
A graceful shutdown cancels tasks and waits for them, and that wait is only as short as the worst task allows. When a task ignores cancellation or the event loop itself is stuck, the process never exits on its own, and the orchestrator eventually kills it without warning. A hard deadline makes the end predictable. Measured on Python 3.14 with 50 well-behaved tasks, each taking 0.2 s to clean up: a normal SIGTERM shutdown exited with code 0 after 0.22 s. Adding one task that caught and ignored CancelledError left the process still running 12 seconds after SIGTERM; an asyncio-level grace period that called os._exit(1) when tasks remained ended it after 2.01 s. Adding instead a task that blocked the loop with a synchronous time.sleep defeated both: the asyncio signal handler never ran, and the process was still running at 12 s — even with a watchdog started from that handler. Starting the watchdog thread from a plain signal.signal handler worked in every case: exit code 70 at 3.00 s. This guide builds a shutdown that ends on time regardless.
Prerequisites¶
- Python 3.11+ on Linux or macOS.
- SIGTERM handling, from handling SIGTERM in asyncio services.
- The topic overview, Graceful Shutdown & Signals.
1. Measure a normal shutdown¶
The baseline: on SIGTERM, set an event; the main coroutine cancels all tasks and waits for them:
async def main():
stop = asyncio.Event()
asyncio.get_running_loop().add_signal_handler(signal.SIGTERM, stop.set)
tasks = [asyncio.create_task(worker(i)) for i in range(50)]
await stop.wait()
for t in tasks:
t.cancel()
await asyncio.gather(*tasks, return_exceptions=True)
asyncio.run(main())
Measured with workers that spend 0.2 s cleaning up: the process exited with code 0 after 0.22 s. Everything about this depends on every task cooperating — receiving the cancellation at an await, finishing its cleanup and re-raising — and on the loop being free to run the signal handler in the first place.
Verify: a SIGTERM test of the service exits with code 0 within the expected cleanup time when nothing is broken.
2. See what breaks it¶
Two kinds of bug defeat a cooperative shutdown. A task that swallows cancellation:
async def swallows_cancel():
while True:
try:
await asyncio.sleep(3600)
except asyncio.CancelledError:
pass # bug: keeps looping after being cancelled
Measured: with one such task among the 50, gather waited forever, asyncio.run never returned, and the process was still running 12 seconds after SIGTERM. A task that blocks the loop with synchronous work:
async def blocks_loop():
await asyncio.sleep(0.5)
time.sleep(30) # bug: the loop cannot run anything for 30 s
Measured: also still running at 12 seconds — and for a different reason. loop.add_signal_handler delivers the signal through the event loop, which was stuck inside time.sleep, so the handler never ran and shutdown never began. Either bug turns a deploy into a wait for the orchestrator's SIGKILL, with no cleanup at all.
Verify: shutdown is tested with a deliberately misbehaving task of each kind, not only with well-behaved ones.
3. Add an asyncio-level grace period¶
The first layer bounds the wait for cooperative tasks and gives up on the rest:
await stop.wait()
for t in tasks:
t.cancel()
done, pending = await asyncio.wait(tasks, timeout=2.0)
if pending:
log.error("%d tasks ignored cancellation: %s", len(pending), [t.get_name() for t in pending])
logging.shutdown()
os._exit(1) # skip waiting for tasks that will never finish
Measured with the cancellation-swallowing task: exit code 1 after 2.01 s, with a log line naming the stuck task — the best possible bug report. os._exit skips atexit handlers and buffered-file flushing, so flush logs before calling it. This layer does nothing for a blocked loop, measured still running at 12 s, because the code that would check the grace period never runs.
Verify: a test with a cancellation-ignoring task exits at the grace period, and the log names the task.
4. Add a watchdog that does not need the loop¶
The last layer must work even when the event loop is stuck. A thread can sleep and call os._exit regardless of what the loop is doing — but it has to be started from somewhere that runs while the loop is blocked. A handler installed with signal.signal runs in the main thread between bytecodes, including during a blocking time.sleep, which is interrupted to run it:
DEADLINE = 3.0
def start_watchdog(deadline: float, code: int = 70):
def run():
time.sleep(deadline)
os.write(2, b"watchdog: shutdown deadline exceeded, exiting\n")
os._exit(code)
threading.Thread(target=run, daemon=True, name="shutdown-watchdog").start()
def install(loop, stop: asyncio.Event):
def on_sigterm(signum, frame):
start_watchdog(DEADLINE) # works even if the loop is blocked
loop.call_soon_threadsafe(stop.set) # graceful path, if the loop can run it
signal.signal(signal.SIGTERM, on_sigterm)
Measured: with the loop-blocking task, the watchdog exited the process with code 70 at 3.00 s; with the cancellation-swallowing task, the same. A watchdog started from loop.add_signal_handler instead never started in the blocked case. In the normal case the process exited cleanly at 0.22 s and the daemon thread died with it. The watchdog writes with os.write rather than logging because the logging lock may be held by the blocked thread.
Verify: with the loop blocked by a synchronous sleep, SIGTERM leads to exit at the watchdog deadline with its exit code.
5. Fit the deadlines to the platform¶
The deadlines are a budget inside the platform's termination grace period — 30 seconds by default in Kubernetes. Leave room at each level:
PLATFORM_GRACE = 30 # terminationGracePeriodSeconds
DRAIN = 20 # stop accepting, finish in-flight requests
CANCEL_GRACE = 5 # cancelled tasks clean up
WATCHDOG = DRAIN + CANCEL_GRACE + 2 # 27 s: before the platform's SIGKILL at 30
Use distinct exit codes — 0 for clean, 1 for the asyncio grace, 70 for the watchdog — so that restarts caused by stuck shutdowns are visible in the platform's records, and alert on the non-zero ones; each is a bug in some task's cancellation handling. For the ordering of draining and pool closing within the cooperative layer, see closing pools cleanly on shutdown and shutting down asyncio pods in Kubernetes.
Verify: the watchdog deadline is shorter than the platform's grace period, and exit codes 1 and 70 raise alerts.
Verification¶
Shutdown is bounded when:
- Cooperative shutdown cancels and drains within its own grace period.
- Tasks that ignore cancellation are logged by name and the process exits at the grace period.
- A watchdog thread started from
signal.signalends the process at a hard deadline, even with a blocked loop. - All deadlines fit inside the platform's grace period, with distinct exit codes.
Diagnostic Hook: when pods take the full termination grace period to stop and then show exit code 137, the process never finished its own shutdown. Run it locally with SIGTERM and a watchdog: exit code 1 at the grace period points to a task swallowing cancellation, and exit code 70 with no grace-period log line points to a blocked event loop.
Pitfalls & edge cases¶
- Tasks that catch
CancelledErrorand continue. Measured: no exit after 12 s. - Blocking calls on the loop. Measured: the asyncio signal handler never ran.
- Starting the watchdog from
loop.add_signal_handler. It never started with a blocked loop. os._exitwithout flushing logs. The reason for the exit is lost.
Frequently Asked Questions¶
Why doesn't my asyncio service exit after SIGTERM?
A task is ignoring cancellation, or the loop is blocked so the handler cannot run. Either kept a test process alive 12 s after SIGTERM.
How do I force an asyncio program to exit after a timeout?
Wait for cancelled tasks with asyncio.wait(..., timeout=grace) and call os._exit if any remain, plus a watchdog thread that calls os._exit at a hard deadline.
Why doesn't loop.add_signal_handler fire when my loop is busy?
It runs the handler on the loop. A blocking time.sleep kept it from running; a handler installed with signal.signal ran in the main thread and started the watchdog.
What exit code should a forced shutdown use?
A distinct non-zero code per layer, such as 1 for the asyncio grace and 70 for the watchdog, so stuck shutdowns show up in platform records.
Related¶
- Graceful Shutdown & Signals — up to the topic overview.
- Shutting down multi-process servers — the same deadlines across worker processes.
- Resilience, Cancellation & Error Handling — the section overview.