Using add_done_callback Without Losing Exceptions¶
add_done_callback is how asyncio itself reacts to completion: gather, wait, TaskGroup and shield are all built on done-callbacks. In application code it is the tool for reacting to a task or future finishing without awaiting it — reporting failures, updating a metric, releasing a slot. It has three behaviours that surprise people. Callbacks never run synchronously, not even when added to a future that is already done; they are scheduled with call_soon. A callback that raises does not affect the future, the awaiter, or the other callbacks — its exception goes only to the loop's exception handler as a log line. And calling result() or exception() on a cancelled future inside a callback raises. In a test with three callbacks where the middle one raised, all three ran in registration order and the only trace of the error was an Exception in callback log record.
Prerequisites¶
- Python 3.11+, stdlib only.
- Futures, from Future Objects & Callbacks.
- Loop exception handler, from installing a custom exception handler.
1. Know when callbacks run¶
import asyncio
async def main() -> None:
loop = asyncio.get_running_loop()
order: list[str] = []
fut = loop.create_future()
fut.add_done_callback(lambda f: order.append("first"))
fut.add_done_callback(lambda f: order.append("second"))
fut.set_result(1)
order.append("after set_result")
await asyncio.sleep(0)
print(order) # ['after set_result', 'first', 'second']
done = loop.create_future()
done.set_result(2)
ran = []
done.add_done_callback(lambda f: ran.append(True))
print(ran) # [] — not run yet, even though the future is done
await asyncio.sleep(0)
print(ran) # [True]
asyncio.run(main())
set_result() schedules every registered callback with call_soon, in registration order, and returns. Adding a callback to a future that is already done schedules it the same way. Either way, callbacks run on a later iteration of the loop, never inside the code that completed the future or registered the callback.
That is a deliberate design: code that calls set_result() cannot be re-entered by arbitrary callbacks mid-statement. The consequence is that you cannot rely on a callback having run immediately after set_result() — if the next line depends on its side effect, the side effect has not happened yet.
Verify: run the snippet; both lists confirm deferred execution.
2. Remember where a raising callback's error goes¶
A done-callback runs as a plain loop callback. If it raises, the loop catches the exception and passes it to the exception handler, which by default logs:
Exception in callback main.<locals>.bad() at app.py:6
handle: <Handle main.<locals>.bad() at app.py:6>
Traceback (most recent call last):
...
ValueError: cb boom
Nothing else happens. The future's result is unchanged, anyone awaiting it gets its result normally, and the remaining callbacks still run — verified with a raising callback between two others, all three ran. That isolation is good for the system and bad for your bug: a callback whose whole job was to record a failure can itself fail, and the only evidence is a log line that does not mention the original future.
Make callbacks defensive, and log with context when they fail:
import functools
import logging
log = logging.getLogger(__name__)
def safe_callback(fn):
@functools.wraps(fn)
def wrapper(fut):
try:
fn(fut)
except Exception:
log.exception("done-callback %s failed for %r", fn.__qualname__, fut)
return wrapper
task.add_done_callback(safe_callback(record_outcome))
Verify: make the wrapped callback raise; the log line names both the callback and the future.
3. Read results safely inside a callback¶
Inside a done-callback the future is done, but "done" has three meanings — result, exception, or cancelled — and the accessors behave differently for each:
| Future state | f.result() |
f.exception() |
f.cancelled() |
|---|---|---|---|
| result set | returns it | None |
False |
| exception set | raises it | returns it | False |
| cancelled | raises CancelledError |
raises CancelledError |
True |
So check cancelled() first, then exception():
def record_outcome(fut: asyncio.Future) -> None:
if fut.cancelled():
metrics.inc("calls", outcome="cancelled")
return
exc = fut.exception() # also marks the exception as retrieved
if exc is not None:
metrics.inc("calls", outcome="error", kind=type(exc).__name__)
return
metrics.inc("calls", outcome="ok")
metrics.observe("payload_bytes", len(fut.result()))
Calling exception() has a side effect worth knowing: it marks the exception as retrieved, so the future will not log "exception was never retrieved" when it is destroyed. If your callback's job is reporting, that is what you want; if another part of the code is responsible for reporting, the callback has silenced it. Handling exceptions in fire-and-forget tasks builds a reporter on exactly this behaviour.
Verify: feed the callback a cancelled future, a failed one and a successful one; none of the three raises.
4. Prefer awaiting when someone is waiting anyway¶
Callbacks are the right tool when nobody awaits the future. When some coroutine already awaits it, put the reaction in that coroutine instead — it runs in a task, can await, has a traceback that leads somewhere, and its exceptions propagate normally:
# callback style: cannot await, errors only logged
def on_done(fut):
asyncio.create_task(store_result(fut.result())) # spawns an untracked task
task.add_done_callback(on_done)
# await style: errors propagate, awaits allowed, ordering is obvious
async def run_and_store():
result = await task
await store_result(result)
The callback version also hides a second problem: create_task inside a callback spawns a task nobody holds a reference to, which can be garbage collected mid-flight and whose failures go unreported. If a callback needs to do async work, have it put the future into a queue that a supervised task drains.
Verify: grep for create_task inside done-callbacks; each is a candidate for the await style or a queue.
5. Remove callbacks that outlive their purpose¶
Callbacks hold references. A long-lived future — a connection's "closed" future, a shutdown event's waiter — with a callback added per request accumulates callbacks, each holding whatever its closure captured:
class Connection:
def __init__(self) -> None:
self.closed = asyncio.get_running_loop().create_future()
async def handle(conn: Connection, request) -> None:
def on_close(_):
request.abort() # captures request
conn.closed.add_done_callback(on_close)
try:
await process(request)
finally:
conn.closed.remove_done_callback(on_close) # otherwise one per request, forever
remove_done_callback(fn) removes every registration of that exact function object and returns how many it removed; because it compares by identity, you must pass the same function object, not an equivalent lambda. This is one of the leak shapes catalogued in leaks from forgotten callbacks and timer handles.
Verify: after 10,000 requests on one connection, len(conn.closed._callbacks) (a private attribute, for inspection only) is zero.
Verification¶
Done-callbacks are used safely when:
- No callback assumes it ran synchronously after
set_result()or after being added. - Every callback checks
cancelled()before reading the result or exception. - Callback failures are logged with context, not just as a bare loop error.
- Callbacks on long-lived futures are removed when their request ends.
Diagnostic Hook: route loop-level "Exception in callback" errors to a dedicated counter through your exception handler, tagged by the callback's qualified name. Any non-zero rate is a bug in reporting or cleanup code — exactly the code that runs least often in tests. In debug mode, slow callbacks over slow_callback_duration are also logged, which catches callbacks doing blocking work.
Pitfalls & edge cases¶
- Expecting synchronous execution. Callbacks are always deferred to a later loop iteration.
- Calling
result()on a cancelled future. RaisesCancelledErrorinside the callback, which is logged as a callback failure. - Removing a different lambda.
remove_done_callbackmatches by identity; an equivalent lambda removes nothing. - Callbacks across threads. Done-callbacks run on the loop thread; a callback added to a
concurrent.futures.Futureruns on whichever thread completes it.
Frequently Asked Questions¶
When does a done-callback run in asyncio?
On a later iteration of the event loop. set_result and set_exception schedule callbacks with call_soon in registration order; adding a callback to a future that is already done also schedules it rather than running it immediately.
What happens if a done-callback raises an exception?
The loop catches it and passes it to the exception handler, which logs "Exception in callback". The future, its awaiters and the other callbacks are unaffected.
Why does task.result() raise CancelledError in my callback?
The task was cancelled. Check task.cancelled() before calling result() or exception(), both of which raise CancelledError for a cancelled future.
Should I use add_done_callback or await?
Await when some coroutine is already waiting for the result: it can do async work and its errors propagate. Use done-callbacks for small synchronous reactions to futures nobody awaits, such as reporting failures of background tasks.
Related¶
- Future Objects & Callbacks — up to the topic overview.
- Chaining futures and transforming results — callbacks used to link one future to another.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.