Bounding Cleanup Time During Cancellation¶
Cancellation in asyncio is cooperative: a cancelled task gets a CancelledError at its current await, runs its except and finally blocks, and only then finishes. Every structure that cancels tasks — asyncio.timeout, TaskGroup, gather, shutdown — waits for that cleanup. So a deadline is only as good as the slowest cleanup it has to wait for. Measured on Python 3.14 with a task inside a TaskGroup inside asyncio.timeout(1.0), whose cleanup awaited a 5-second close: the caller received TimeoutError at 6.01 s, not 1 s. Wrapping the cleanup in its own asyncio.timeout(0.5) made the caller's timeout arrive at 1.5 s. Moving the slow close into a detached task let the caller continue at 1.0 s while the close finished in the background at 6.01 s. And cancelling the task a second time during cleanup ended it at 1.2 s — with its cleanup never completed. This guide bounds cleanup so that deadlines hold.
Prerequisites¶
- Python 3.11+ for
asyncio.timeoutandTaskGroup. - Cancellation basics, from cancelling a task and waiting for it to finish.
- The topic overview, Cancellation Patterns.
1. See a timeout stretched by cleanup¶
A worker that closes a resource gracefully when cancelled:
async def worker(conn):
try:
await conn.stream_forever()
except asyncio.CancelledError:
await conn.close_gracefully() # 5 s when the peer is unresponsive
raise
async with asyncio.timeout(1.0):
async with asyncio.TaskGroup() as tg:
tg.create_task(worker(conn))
Measured: at 1.0 s the timeout cancelled the TaskGroup's host, the group cancelled the worker, and the worker began its 5-second close. The group waited, so the timeout waited, and the caller received TimeoutError at 6.01 s. Nothing was wrong with the timeout; it delivered cancellation on time. The cancelled code decided how long it took to finish. The same stretch applies to request deadlines, shutdown deadlines and any TaskGroup whose sibling failed.
Verify: for each operation run under a deadline, the worst-case cleanup time of everything it might cancel is known.
2. Put a timeout on the cleanup itself¶
A cancelled task can still await, and it can still use asyncio.timeout. Bound the cleanup explicitly and fall back to an abrupt close:
async def worker(conn):
try:
await conn.stream_forever()
except asyncio.CancelledError:
try:
async with asyncio.timeout(0.5):
await conn.close_gracefully()
except TimeoutError:
conn.abort() # immediate: drop the socket, no handshake
raise # always re-raise the cancellation
Measured: the cleanup gave up after 0.5 s, and the caller's TimeoutError arrived at 1.5 s — the 1-second deadline plus the cleanup bound. The raise at the end matters: catching CancelledError without re-raising turns a cancellation into normal completion, and the TaskGroup or timeout above would never learn the task stopped. The abrupt fallback must not await anything unbounded itself.
Verify: every except CancelledError or finally block that awaits has a timeout around the await, and re-raises.
3. Detach cleanup that must finish¶
Some cleanup should complete even though the caller cannot wait for it — flushing a buffer to storage, releasing a distributed lock. Hand it to a separate task that the cancellation does not reach:
background: set[asyncio.Task] = set()
def detach(coro):
task = asyncio.create_task(coro)
background.add(task) # keep a reference until it finishes
task.add_done_callback(background.discard)
async def worker(conn):
try:
await conn.stream_forever()
except asyncio.CancelledError:
detach(conn.close_gracefully())
raise
Measured: the caller received TimeoutError at 1.0 s, and the detached close finished at 6.01 s in the background. The trade is visibility: the caller has moved on, so the detached work needs its own error handling and logging, and a reference — the event loop holds only weak references to tasks, and an unreferenced task can be garbage-collected before it finishes. At process shutdown, wait for the background set with its own deadline; otherwise asyncio.run cancels whatever is still running.
Verify: detached cleanup tasks are referenced until done, log their failures, and are awaited with a deadline at shutdown.
4. Do not cancel twice to hurry things up¶
When cleanup is slow, it is tempting to cancel again. Measured: cancelling the task a second time at 1.2 s, during its 5-second close, interrupted the close at that await, and the task ended at 1.2 s — with the close never completed and no log line from the cleanup. The second cancellation landed inside the cleanup code, which is the worst place for it. Repeated cancellation is how resources leak during shutdown:
task.cancel()
try:
async with asyncio.timeout(grace):
await task # let cleanup run within the grace period
except TimeoutError:
task.cancel() # last resort: interrupts cleanup, may leak
log.warning("cleanup of %r exceeded %.1fs; forced", task, grace)
Escalate only after a bounded wait, log it, and make the code being interrupted able to tolerate it — for example, with the abrupt fallback from step 2 in a finally block, so even an interrupted cleanup drops the socket. For shutdown across many tasks, the same pattern is applied per process in enforcing a hard shutdown deadline.
Verify: code that cancels a task twice does so only after waiting a bounded grace period, and logs it.
5. Budget cleanup into the deadline¶
If an operation must finish within a deadline including cleanup, cancel early enough to leave room for it. Split the budget explicitly:
async def call_with_budget(op, total: float, cleanup_budget: float = 0.5):
async with asyncio.timeout(total - cleanup_budget): # work gets the rest
return await op()
# each cancellable resource bounds its own cleanup to cleanup_budget (step 2)
By construction, with a 1-second total and a 0.5-second cleanup budget, work is cancelled at 0.5 s and cleanup — bounded at 0.5 s — finishes by 1.0 s, so the caller's promise to its own caller holds. Document the cleanup budget next to the resources that use it, since a library upgrade that makes a close slower silently breaks the arithmetic. For related shutdown work on executors and generators, see shutting down async generators and executors cleanly.
Verify: an end-to-end test with the slowest possible cleanup completes within the outer deadline.
Verification¶
Cleanup is bounded when:
- Every await in a cancellation path has a timeout, with an abrupt fallback.
CancelledErroris always re-raised after cleanup.- Cleanup that must finish is detached, referenced, and awaited at shutdown.
- Second cancellations happen only after a bounded wait, and are logged.
Diagnostic Hook: when an operation under asyncio.timeout(1) regularly takes far longer, time the cleanup of what it cancels. A 5-second graceful close turned a 1-second timeout into 6.01 seconds in this test.
Pitfalls & edge cases¶
- Unbounded awaits in
except CancelledError. Measured: a 1 s timeout took 6.01 s. - Swallowing
CancelledErrorin cleanup. The canceller never learns the task stopped. - Unreferenced detached tasks. They can be garbage-collected before finishing.
- Cancelling again to speed up. Measured: cleanup interrupted at 1.2 s.
Frequently Asked Questions¶
Why does asyncio.timeout take longer than the timeout?
It waits for the cancelled code to finish. A task whose cleanup awaited a 5 s close turned a 1 s timeout into 6.01 s.
Can I await inside an except CancelledError block?
Yes, and you can use asyncio.timeout there too: a 0.5 s bound on cleanup made the caller's timeout arrive at 1.5 s. Always re-raise afterwards.
What happens if a task is cancelled during its cleanup?
The cleanup is interrupted at its current await. Cancelling again at 1.2 s ended the task with its 5 s close never completed.
How do I make cleanup finish without blocking the caller?
Start it as a separate, referenced task. The caller continued at 1.0 s and the close finished in the background at 6.01 s.
Related¶
- Cancellation Patterns — up to the topic overview.
- Comparing gather and TaskGroup cancellation — both wait for this cleanup.
- Resilience, Cancellation & Error Handling — the section overview.