Skip to content

Using add_done_callback Without Losing Exceptions

add_done_callback is how asyncio itself reacts to completion: gather, wait, TaskGroup and shield are all built on done-callbacks. In application code it is the tool for reacting to a task or future finishing without awaiting it — reporting failures, updating a metric, releasing a slot. It has three behaviours that surprise people. Callbacks never run synchronously, not even when added to a future that is already done; they are scheduled with call_soon. A callback that raises does not affect the future, the awaiter, or the other callbacks — its exception goes only to the loop's exception handler as a log line. And calling result() or exception() on a cancelled future inside a callback raises. In a test with three callbacks where the middle one raised, all three ran in registration order and the only trace of the error was an Exception in callback log record.

Prerequisites

1. Know when callbacks run

import asyncio


async def main() -> None:
    loop = asyncio.get_running_loop()
    order: list[str] = []
    fut = loop.create_future()
    fut.add_done_callback(lambda f: order.append("first"))
    fut.add_done_callback(lambda f: order.append("second"))
    fut.set_result(1)
    order.append("after set_result")
    await asyncio.sleep(0)
    print(order)                     # ['after set_result', 'first', 'second']

    done = loop.create_future()
    done.set_result(2)
    ran = []
    done.add_done_callback(lambda f: ran.append(True))
    print(ran)                       # [] — not run yet, even though the future is done
    await asyncio.sleep(0)
    print(ran)                       # [True]


asyncio.run(main())

set_result() schedules every registered callback with call_soon, in registration order, and returns. Adding a callback to a future that is already done schedules it the same way. Either way, callbacks run on a later iteration of the loop, never inside the code that completed the future or registered the callback.

That is a deliberate design: code that calls set_result() cannot be re-entered by arbitrary callbacks mid-statement. The consequence is that you cannot rely on a callback having run immediately after set_result() — if the next line depends on its side effect, the side effect has not happened yet.

Verify: run the snippet; both lists confirm deferred execution.

set_result schedules callbacks; it never runs them A sequence of 5 messages between 4 participants. set_result schedules callbacks; it never runs them completing code future ready queue loop set_result(1) call_soon(first), call_soon(second) continues: callbacks have not run next iteration: pop first pop second Callbacks always run later, in registration order, on the loop's thread.

2. Remember where a raising callback's error goes

A done-callback runs as a plain loop callback. If it raises, the loop catches the exception and passes it to the exception handler, which by default logs:

Exception in callback main.<locals>.bad() at app.py:6
handle: <Handle main.<locals>.bad() at app.py:6>
Traceback (most recent call last):
  ...
ValueError: cb boom

Nothing else happens. The future's result is unchanged, anyone awaiting it gets its result normally, and the remaining callbacks still run — verified with a raising callback between two others, all three ran. That isolation is good for the system and bad for your bug: a callback whose whole job was to record a failure can itself fail, and the only evidence is a log line that does not mention the original future.

Make callbacks defensive, and log with context when they fail:

import functools
import logging

log = logging.getLogger(__name__)


def safe_callback(fn):
    @functools.wraps(fn)
    def wrapper(fut):
        try:
            fn(fut)
        except Exception:
            log.exception("done-callback %s failed for %r", fn.__qualname__, fut)
    return wrapper


task.add_done_callback(safe_callback(record_outcome))

Verify: make the wrapped callback raise; the log line names both the callback and the future.

3. Read results safely inside a callback

Inside a done-callback the future is done, but "done" has three meanings — result, exception, or cancelled — and the accessors behave differently for each:

Future state f.result() f.exception() f.cancelled()
result set returns it None False
exception set raises it returns it False
cancelled raises CancelledError raises CancelledError True

So check cancelled() first, then exception():

def record_outcome(fut: asyncio.Future) -> None:
    if fut.cancelled():
        metrics.inc("calls", outcome="cancelled")
        return
    exc = fut.exception()                      # also marks the exception as retrieved
    if exc is not None:
        metrics.inc("calls", outcome="error", kind=type(exc).__name__)
        return
    metrics.inc("calls", outcome="ok")
    metrics.observe("payload_bytes", len(fut.result()))

Calling exception() has a side effect worth knowing: it marks the exception as retrieved, so the future will not log "exception was never retrieved" when it is destroyed. If your callback's job is reporting, that is what you want; if another part of the code is responsible for reporting, the callback has silenced it. Handling exceptions in fire-and-forget tasks builds a reporter on exactly this behaviour.

Verify: feed the callback a cancelled future, a failed one and a successful one; none of the three raises.

Reading a finished future inside a callback A decision on Is fut.cancelled() true with 3 outcomes. Reading a finished future inside a callback Is fut.cancelled() true? yes stop here result() would raise no, exception() is set handle the error marks it retrieved no, exception() is None fut.result() is safe the value The order of checks is what keeps a reporting callback from becoming a new error.

4. Prefer awaiting when someone is waiting anyway

Callbacks are the right tool when nobody awaits the future. When some coroutine already awaits it, put the reaction in that coroutine instead — it runs in a task, can await, has a traceback that leads somewhere, and its exceptions propagate normally:

# callback style: cannot await, errors only logged
def on_done(fut):
    asyncio.create_task(store_result(fut.result()))   # spawns an untracked task

task.add_done_callback(on_done)


# await style: errors propagate, awaits allowed, ordering is obvious
async def run_and_store():
    result = await task
    await store_result(result)

The callback version also hides a second problem: create_task inside a callback spawns a task nobody holds a reference to, which can be garbage collected mid-flight and whose failures go unreported. If a callback needs to do async work, have it put the future into a queue that a supervised task drains.

Verify: grep for create_task inside done-callbacks; each is a candidate for the await style or a queue.

5. Remove callbacks that outlive their purpose

Callbacks hold references. A long-lived future — a connection's "closed" future, a shutdown event's waiter — with a callback added per request accumulates callbacks, each holding whatever its closure captured:

class Connection:
    def __init__(self) -> None:
        self.closed = asyncio.get_running_loop().create_future()


async def handle(conn: Connection, request) -> None:
    def on_close(_):
        request.abort()                        # captures request

    conn.closed.add_done_callback(on_close)
    try:
        await process(request)
    finally:
        conn.closed.remove_done_callback(on_close)   # otherwise one per request, forever

remove_done_callback(fn) removes every registration of that exact function object and returns how many it removed; because it compares by identity, you must pass the same function object, not an equivalent lambda. This is one of the leak shapes catalogued in leaks from forgotten callbacks and timer handles.

Verify: after 10,000 requests on one connection, len(conn.closed._callbacks) (a private attribute, for inspection only) is zero.

Good and bad uses of add_done_callback A grid of 5 rows by 3 columns. Good and bad uses of add_done_callback use fit why report a background failure good nobody else awaits it release a slot or reference good synchronous, must always run record a metric good cheap and synchronous start more async work poor spawns untracked tasks one per request on a shared future leak remove it in finally Callbacks are for small synchronous reactions; async follow-up work belongs in a task.

Verification

Done-callbacks are used safely when:

  • No callback assumes it ran synchronously after set_result() or after being added.
  • Every callback checks cancelled() before reading the result or exception.
  • Callback failures are logged with context, not just as a bare loop error.
  • Callbacks on long-lived futures are removed when their request ends.

Diagnostic Hook: route loop-level "Exception in callback" errors to a dedicated counter through your exception handler, tagged by the callback's qualified name. Any non-zero rate is a bug in reporting or cleanup code — exactly the code that runs least often in tests. In debug mode, slow callbacks over slow_callback_duration are also logged, which catches callbacks doing blocking work.

Pitfalls & edge cases

  • Expecting synchronous execution. Callbacks are always deferred to a later loop iteration.
  • Calling result() on a cancelled future. Raises CancelledError inside the callback, which is logged as a callback failure.
  • Removing a different lambda. remove_done_callback matches by identity; an equivalent lambda removes nothing.
  • Callbacks across threads. Done-callbacks run on the loop thread; a callback added to a concurrent.futures.Future runs on whichever thread completes it.

Frequently Asked Questions

When does a done-callback run in asyncio?

On a later iteration of the event loop. set_result and set_exception schedule callbacks with call_soon in registration order; adding a callback to a future that is already done also schedules it rather than running it immediately.

What happens if a done-callback raises an exception?

The loop catches it and passes it to the exception handler, which logs "Exception in callback". The future, its awaiters and the other callbacks are unaffected.

Why does task.result() raise CancelledError in my callback?

The task was cancelled. Check task.cancelled() before calling result() or exception(), both of which raise CancelledError for a cancelled future.

Should I use add_done_callback or await?

Await when some coroutine is already waiting for the result: it can do async work and its errors propagate. Use done-callbacks for small synchronous reactions to futures nobody awaits, such as reporting failures of background tasks.