Handling Exceptions in Fire-and-Forget Tasks¶
A fire-and-forget task is one nobody awaits: asyncio.create_task(send_audit_event(...)) and move on. When it fails, its exception is stored on the task object, and asyncio reports it only if the task is destroyed without anyone having looked. When that happens depends entirely on who holds a reference. In a test with three identical failing tasks: the one with no reference was logged within the same loop iteration; the one kept in a set with a discard done-callback was logged as soon as the callback dropped it; and the one kept in a set without that callback was not logged until the interpreter exited — after the program had finished and, in a real service, hours or days after the failure. This guide makes background failures visible at the moment they happen.
Prerequisites¶
- Python 3.11+, stdlib only; behaviour below was checked on Python 3.14.
- Why tasks need strong references, from preventing task garbage collection.
- The loop exception handler, from installing a custom exception handler.
1. Reproduce the three behaviours¶
Run this once and read the order of the output; it is the whole problem in miniature:
import asyncio
import logging
logging.basicConfig(level=logging.ERROR, format="%(levelname)s %(message)s")
keep: set[asyncio.Task] = set()
async def boom() -> None:
raise ValueError("boom")
async def main() -> None:
asyncio.create_task(boom()) # A: no reference
await asyncio.sleep(0.01)
t = asyncio.create_task(boom()) # B: referenced forever
keep.add(t)
await asyncio.sleep(0.01)
t = asyncio.create_task(boom()) # C: referenced until done
keep.add(t)
t.add_done_callback(keep.discard)
await asyncio.sleep(0.01)
print("main finished")
asyncio.run(main())
print("interpreter shutting down")
Task A and task C each produce ERROR Task exception was never retrieved before main finished. Task B's error appears only after interpreter shutting down. The report is emitted from the task's finaliser, so it fires when the last reference disappears — immediately for A (CPython's reference counting frees it at once), on the done-callback for C, and at process exit for B.
The trap: B is what most people write after learning that unreferenced tasks can be garbage collected. Holding the reference fixes the GC problem and creates a silent-failure problem.
Verify: your output should show two errors before main finished and one after interpreter shutting down.
2. Report at completion with a done-callback¶
Never rely on the finaliser. Attach a done-callback that retrieves the exception and logs it with context the moment the task finishes:
import asyncio
import logging
log = logging.getLogger("tasks")
def _report(task: asyncio.Task) -> None:
if task.cancelled():
log.debug("task %s cancelled", task.get_name())
return
exc = task.exception() # marks the exception as retrieved
if exc is not None:
log.error("background task %s failed", task.get_name(),
exc_info=(type(exc), exc, exc.__traceback__))
Calling task.exception() marks the exception as retrieved, so the finaliser will not log it a second time. Passing the explicit exc_info tuple keeps the full traceback: the traceback was captured when the coroutine raised, and the log record prints it as if it were being handled now.
Check cancelled() first. Calling exception() on a cancelled task raises CancelledError — inside a done-callback, that becomes a loop-level "Exception in callback" error, which is noise in the best case.
Verify: a failing task now produces exactly one log line, named, with the original traceback, as soon as it fails.
3. Wrap it in one spawn helper¶
Spread across a codebase, the reference set and the done-callback get forgotten in exactly the places that matter. Put them in one function and ban bare create_task for background work:
import asyncio
from collections.abc import Coroutine
from typing import Any
_background: set[asyncio.Task] = set()
def spawn(coro: Coroutine[Any, Any, Any], *, name: str) -> asyncio.Task:
"""Start supervised background work: referenced, named, and reported on failure."""
task = asyncio.create_task(coro, name=name)
_background.add(task)
task.add_done_callback(_background.discard)
task.add_done_callback(_report)
return task
async def cancel_background(timeout: float = 5.0) -> None:
"""Call from shutdown: cancel everything still running and wait for it."""
tasks = list(_background)
for t in tasks:
t.cancel()
async with asyncio.timeout(timeout):
await asyncio.gather(*tasks, return_exceptions=True)
The name argument is mandatory on purpose: an error log that says Task-8812 failed is close to useless, while audit:order-5531 failed points at the caller. The shutdown helper is the other half of supervision — a background set that nobody cancels leaves tasks that are destroyed while pending, which asyncio logs as Task was destroyed but it is pending!.
A lint rule makes the ban stick: forbid asyncio.create_task outside this module with a custom flake8/ruff rule or a grep in CI.
Verify: _background is empty after every task finishes; shutdown cancels the remainder within the timeout.
4. Decide whether failure should stop anything¶
Logging is the floor, not the ceiling. Some background work is optional — a cache warm, an audit event that is also written elsewhere — and logging the failure is all that is needed. Other work is load-bearing: a consumer loop, a lease renewer, a heartbeat. If that dies, the service is broken while still answering health checks.
For load-bearing work, do not fire and forget. Put it under a TaskGroup owned by the service's main coroutine, so its failure tears the service down and the orchestrator restarts it:
async def serve() -> None:
async with asyncio.TaskGroup() as tg:
tg.create_task(consume_orders(), name="consumer")
tg.create_task(renew_lease(), name="lease")
tg.create_task(run_http_server(), name="http")
# if any of the three raises, the others are cancelled and the error propagates
The middle ground is a supervisor that restarts a failed task with backoff, covered in restarting crashed workers with a supervisor task. Choose deliberately; the default of "log and carry on" is wrong for anything the service cannot work without.
Verify: kill the consumer loop with an injected exception and confirm the process exits non-zero instead of carrying on without it.
5. Catch the stragglers at the loop level¶
Some background work is started by libraries, not by you. Their failures reach the loop's exception handler through the finaliser path, with a context dict containing message, exception and future. Route them to the same place as everything else:
def loop_exception_handler(loop: asyncio.AbstractEventLoop, context: dict) -> None:
exc = context.get("exception")
task = context.get("future") or context.get("task")
name = task.get_name() if isinstance(task, asyncio.Task) else "-"
log.error("unhandled in loop: %s (task=%s)", context.get("message"), name,
exc_info=(type(exc), exc, exc.__traceback__) if exc else None)
async def main() -> None:
asyncio.get_running_loop().set_exception_handler(loop_exception_handler)
Running tests with -W error::RuntimeWarning and PYTHONASYNCIODEBUG=1 adds the creation traceback of each task to these messages, which tells you which library created the orphan.
Verify: a failing task created by third-party code produces a log line through your handler, not a bare stderr print.
Verification¶
Background failures are handled when:
- Every failure is logged within one loop iteration of happening, with the task's name and the original traceback.
- No
Task exception was never retrievedlines appear at interpreter exit in a test run that exercises failures. - No
Task was destroyed but it is pending!lines appear on shutdown. - Load-bearing tasks crash the service when they die, rather than leaving a zombie process.
Diagnostic Hook: count background failures by task-name prefix (audit:, refresh:) as a metric, and expose len(_background) as a gauge. A rising failure rate is the alert; a background set that only grows is the leak that tells you something is spawning faster than it finishes.
Pitfalls & edge cases¶
exception()on a cancelled task raises. Always checkcancelled()first in a done-callback.- Swallowing
CancelledErrorinside the task body. The task then completes "successfully" and nothing reports that it was interrupted. - Logging with
log.exception()in a done-callback. There is no active exception there, so the traceback is lost; passexc_infoexplicitly. - Done-callbacks that raise. They are run via
call_soon, so the error goes to the loop handler and the remaining callbacks still run — but your reporting may be the thing that failed. gather(*tasks)withoutreturn_exceptions=Trueat shutdown. The first failure propagates and the remaining tasks are left un-awaited.
Frequently Asked Questions¶
Why is 'Task exception was never retrieved' logged so late?
The message is emitted from the task's finaliser, when the task object is destroyed. If you hold the task in a set that never releases it, that only happens at interpreter shutdown. Retrieve the exception in a done-callback instead.
How do I log exceptions from asyncio.create_task?
Add a done-callback that checks task.cancelled(), then calls task.exception() and logs it with exc_info set to the exception's type, value and traceback. Wrap create_task, the reference set and the callback in one spawn() helper so it is never forgotten.
Should I use a TaskGroup instead of fire-and-forget tasks?
For work the service cannot run without, yes: a TaskGroup in the main coroutine makes a failure stop the service so it can be restarted. For optional work that should not block the caller, a supervised spawn helper is the better fit.
Does task.exception() clear the error?
It marks it as retrieved, so the finaliser will not log it again. The exception stays on the task and later calls return the same object.
How do I find which code created an orphaned failing task?
Run with PYTHONASYNCIODEBUG=1. Debug mode records where each task was created and includes that traceback in the loop's error messages.
Related¶
- Task Scheduling & Lifecycle — up to the topic overview.
- Naming and tracking tasks for observability — making the names in these logs meaningful.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.