Skip to content

Reporting Background Task Errors in asyncio

asyncio reports an exception from a task nobody awaited with the log message "Task exception was never retrieved" — but only when the task object is destroyed. Whether and when that happens depends on what still references the task, so background failures can be logged at once, much later, or not at all. Measured on Python 3.14 with a background task that raised after 10 ms: with no reference kept, the error was logged 21 ms after creation. With the task kept in a list, it was not logged within 2 seconds — never, while the list lived. With the common pattern of a module-level set and add_done_callback(set.discard), it was logged after 21 ms when the task was spawned from a helper function, and not at all when the calling function still held the task in a local variable. A task caught in a reference cycle was reported only when the garbage collector ran. A done callback that logged task.exception() reported every failure 21 ms after creation, whatever held the task. This guide makes background failures visible deterministically.

Prerequisites

1. See when asyncio reports an unretrieved exception

asyncio logs an unretrieved task exception from the task's finalizer. The timing therefore follows Python's memory management, not the failure:

async def boom():
    await asyncio.sleep(0.01)
    raise RuntimeError("background job failed")

asyncio.create_task(boom())                  # nothing keeps a reference

Measured: logged 21 ms after creation — the task finished at 10 ms and was freed as soon as the loop dropped its last reference. Keep a reference, and the report waits for that reference to go: with the task appended to a list that lived for the rest of the test, nothing was logged within 2 seconds. Put the task in a reference cycle — for example, an object that references itself and holds the task — and the report waited for the cyclic garbage collector: none before gc.collect(), reported right after it. The failure happened at 10 ms in every case; only its visibility moved.

Verify: in a test, make a background task fail and check how long its error takes to appear in the logs.

A background task raising after 10 ms A grid of 6 rows by 2 columns. A background task raising after 10 ms how the task is held error appears in logs no reference after 21 ms (finalizer) kept in a list not within 2 s set + discard on done, spawned via helper after 21 ms set + discard, caller keeps a local variable not within 1 s in a reference cycle only when gc runs done callback logs task.exception() after 21 ms, always Python 3.14; 'never retrieved' is logged from the task's finalizer.

2. Keep a reference, and know what it costs

The event loop holds only weak references to tasks, so a task with no other reference can be garbage-collected while it is still running. The standard advice is to keep tasks in a set and discard them when done:

background: set[asyncio.Task] = set()

def spawn(coro) -> None:
    task = asyncio.create_task(coro)
    background.add(task)
    task.add_done_callback(background.discard)

Measured: when spawned through this helper, the failing task was reported after 21 ms — discarding it from the set dropped its last reference. When a calling function kept its own variable pointing at the task instead, the report did not appear within 1 second, because that local reference kept the task alive. The pattern protects running tasks from collection; it does not guarantee their failures are reported, because any other reference — a variable, a dictionary of jobs, a debugging list — delays the finalizer indefinitely.

Verify: no code path keeps references to fire-and-forget tasks beyond the background set.

3. Report failures from a done callback

Reporting should not depend on garbage collection. Attach a callback that retrieves the exception and logs it as soon as the task finishes:

def _report(task: asyncio.Task) -> None:
    if task.cancelled():
        return
    exc = task.exception()                           # retrieving it suppresses "never retrieved"
    if exc is not None:
        log.error("background task %s failed", task.get_name(), exc_info=exc)
        background_failures.labels(task=task.get_name()).inc()

def spawn(coro, name: str) -> asyncio.Task:
    task = asyncio.create_task(coro, name=name)
    background.add(task)
    task.add_done_callback(background.discard)
    task.add_done_callback(_report)
    return task

Measured: the failure was logged 21 ms after creation — as soon as the task finished — even with the task still referenced from a list. Calling task.exception() marks the exception as retrieved, so the finalizer does not log it a second time. Name tasks: the log line and the metric label then say which job failed, not just that one did.

Verify: every fire-and-forget task is created through the helper, and a test shows its failure logged within milliseconds while a reference is still held.

Spawning a background task that reports its own failure A flow of 5 stages. Spawning a background task that reports its own failure create_task(name=...) named for logs and metrics background.add not collected while running done callbacks discard + report Task fails report: task.exception(), log, metric Discard last reference released Measured: reported 21 ms after creation regardless of other references.

4. Catch the rest with a loop exception handler

Tasks created by libraries do not go through your helper. A loop-level exception handler sees everything asyncio itself reports — including "Task exception was never retrieved" — and can route it into structured logs and metrics:

def handle_loop_exception(loop, context):
    exc = context.get("exception")
    log.error("asyncio: %s", context.get("message"), exc_info=exc,
              extra={"task": repr(context.get("task") or context.get("future"))})
    unhandled_async_errors.inc()

asyncio.get_running_loop().set_exception_handler(handle_loop_exception)

Measured: for a failing task spawned through the set-and-discard helper, the handler received "Task exception was never retrieved" with the RuntimeError attached, 10 ms after creation. It still depends on the finalizer — it fires when asyncio would have logged — so it complements the done callback rather than replacing it. It also receives errors from callbacks scheduled with call_soon, which have no task to report through.

Verify: the loop's exception handler is installed at start-up and feeds a metric that alerts above zero.

5. Prefer structure where you can

Fire-and-forget tasks need reporting because nothing awaits them. Where the work has a natural owner, give it one: a TaskGroup or a supervisor task awaits its children and receives their exceptions directly, with no dependence on callbacks or garbage collection.

async def background_services(app):
    async with asyncio.TaskGroup() as tg:                  # one owner for the long-running loops
        tg.create_task(refresh_cache_forever(app))
        tg.create_task(flush_metrics_forever(app))

@asynccontextmanager
async def lifespan(app):
    owner = asyncio.create_task(background_services(app), name="background-services")
    owner.add_done_callback(_report)                      # a failing loop is reported at once
    yield
    owner.cancel()
    await asyncio.gather(owner, return_exceptions=True)

Use the reporting helper for genuinely detached work — a write-behind, a notification — and structured ownership for long-running loops, which should also restart on failure, as in restarting crashed workers with a supervisor task. Either way, a background failure is logged when it happens, with the task's name, and counted.

Verify: long-running background loops are owned by a TaskGroup or supervisor, and detached tasks go through the reporting helper.

Delay before a 10 ms failure was reported 5 horizontal bars comparing done callback logs task.exception() with the others. Delay before a 10 ms failure was reported done callback logs task.exception() 21 ms no reference (finalizer) 21 ms set + discard via helper (finalizer) 21 ms set + discard, caller keeps a local > 1 s kept in a list > 2 s The last two were still unreported when the test stopped watching.

Verification

Background failures are reported when:

  • Every detached task is created through a helper that adds a reporting done callback.
  • Reports name the task and increment a metric.
  • A loop exception handler routes asyncio's own reports into structured logs.
  • Long-running loops have an owner — a TaskGroup or a supervisor — instead of being detached.

Diagnostic Hook: when a background job silently stops working and the logs show nothing, look for references to its task outside the background set. Kept in a list, a failed task's error was not logged within 2 seconds in this test; a done callback reported it after 21 ms.

Pitfalls & edge cases

  • Relying on "Task exception was never retrieved". It appears only when the task is freed.
  • Extra references to background tasks. Measured: no report while a list held the task.
  • Reference cycles. Measured: reported only when the garbage collector ran.
  • Unnamed tasks. Reports say a task failed, not which job it was.

Frequently Asked Questions

Why don't I see errors from my asyncio background tasks?

asyncio logs them only when the task object is destroyed. While anything references it — a list, a variable — nothing is logged; kept in a list, nothing appeared within 2 s.

When is 'Task exception was never retrieved' logged?

From the task's finalizer: 21 ms after creation for an unreferenced task here, and only when the garbage collector ran for a task in a reference cycle.

How do I log exceptions from fire-and-forget tasks reliably?

Add a done callback that calls task.exception() and logs it with the task name. It reported the failure 21 ms after creation regardless of references.

Does keeping tasks in a set hide their errors?

The set alone does not if done tasks are discarded and nothing else holds them. Any other reference delays the report indefinitely.