Skip to content

Adopting TaskGroup and asyncio.timeout When Upgrading from 3.10

Python 3.11 is the release that made structured concurrency part of asyncio, and raising a codebase's minimum version from 3.10 to 3.11 is the moment to adopt it. The rewrite is not mechanical, because the new primitives deliberately behave differently from the ones they replace. A side-by-side test made the difference concrete: with one failing task and three slow siblings, asyncio.gather() raised the failure at 0.05 s and let all three siblings finish in the background; a TaskGroup raised an ExceptionGroup at the same moment and cancelled all three — none finished. That is the improvement, and it is also a behaviour change at every call site. This guide converts gather, wait_for, async-timeout and error handling in an order that keeps each step reviewable.

Prerequisites

1. Inventory the call sites by intent

gather() calls were written with one of three intents, and each maps to a different 3.11 construct:

grep -rn "asyncio.gather(" --include=*.py src | wc -l
grep -rn "return_exceptions=True" --include=*.py src
grep -rn "wait_for(\|async_timeout\|async-timeout" --include=*.py src
  • "All of these must succeed" — plain gather(*coros) whose caller treats any exception as failure. → TaskGroup.
  • "Run all, collect every outcome" — gather(..., return_exceptions=True) followed by inspecting results. → keep gather, or a TaskGroup whose children catch their own errors.
  • "Bound the time" — wait_for(coro, t) or async with async_timeout.timeout(t). → asyncio.timeout().

Classify before rewriting: a gather whose caller did rely on siblings finishing after a failure — rare, but it happens with best-effort writes — must not become a TaskGroup.

Verify: each call site has an intent label in the review checklist.

One failure, three slow siblings 4 lanes over time. One failure, three slow siblings failing task raises at 50 ms gather siblings running keep running unobserved TaskGroup siblings running cancelled caller sees the error time → Measured: both raise at 0.05 s; only gather's siblings went on to finish, unobserved.

2. Convert "all must succeed" gathers to TaskGroup

# before (3.10)
user, orders, prefs = await asyncio.gather(
    load_user(uid), load_orders(uid), load_prefs(uid)
)

# after (3.11+)
async with asyncio.TaskGroup() as tg:
    user_t = tg.create_task(load_user(uid))
    orders_t = tg.create_task(load_orders(uid))
    prefs_t = tg.create_task(load_prefs(uid))
user, orders, prefs = user_t.result(), orders_t.result(), prefs_t.result()

What changes for callers: a failure now arrives as an ExceptionGroup containing every child failure, not as the first exception. Code that wrapped the old gather in except NotFound: will no longer catch it — the group is not a NotFound. Update the handler in the same change, with except*:

async def get_profile(uid):
    missing = False
    try:
        async with asyncio.TaskGroup() as tg:
            user_t = tg.create_task(load_user(uid))
            orders_t = tg.create_task(load_orders(uid))
    except* NotFound:
        missing = True                         # return/break/continue are a SyntaxError in except*
    except* (TimeoutError, ConnectionError) as eg:
        raise ServiceUnavailable() from eg
    if missing:
        return None
    return build_profile(user_t.result(), orders_t.result())

except* is syntax, which is why this step needs the 3.11 floor rather than a backport. More patterns, including partial results, are in migrating from gather to TaskGroup.

Verify: the tests that covered the old error paths still pass after updating their expected exception types.

3. Keep gather where you want every outcome

gather(..., return_exceptions=True) has no direct TaskGroup equivalent, and does not need one: it is the right tool when every outcome matters and no failure should stop the others.

results = await asyncio.gather(*(notify(u) for u in users), return_exceptions=True)
failed = [(u, r) for u, r in zip(users, results) if isinstance(r, BaseException)]

If you prefer TaskGroup everywhere for consistency, push the error handling into the children, so the group only sees failures that should abort:

async def notify_safely(u):
    try:
        await notify(u)
    except NotificationError as exc:      # expected, per-item
        failures.append((u, exc))


async with asyncio.TaskGroup() as tg:
    for u in users:
        tg.create_task(notify_safely(u))

One correction to a common belief: return_exceptions=True also returns CancelledError instances as results if a child is cancelled, rather than propagating cancellation. Check for BaseException, not Exception, when filtering.

Verify: a run with injected per-item failures completes all items and reports exactly the failed ones.

4. Replace wait_for and async-timeout with asyncio.timeout

# before
result = await asyncio.wait_for(fetch(url), timeout=5)

async with async_timeout.timeout(5):          # pip install async-timeout
    result = await fetch(url)

# after
async with asyncio.timeout(5):
    result = await fetch(url)

asyncio.timeout() wraps a block, so several awaits share one deadline — the usual intent for a request budget. It raises TimeoutError (the builtin; asyncio.TimeoutError is an alias since 3.11), and its deadline can be moved with reschedule(), which wait_for never supported. On 3.12+, wait_for itself is implemented with timeout() and runs in the caller's task; on 3.10 and 3.11 it created a separate task, so context-variable changes inside it were invisible to the caller — a difference the version probe measured directly. Converting to timeout() removes that version dependence. Deadline handling in depth is in choosing asyncio.timeout vs wait_for.

Drop the async-timeout dependency once no call site imports it.

Verify: grep -rn "async_timeout" returns nothing, and timeout tests still observe TimeoutError.

3.10 constructs and their 3.11 replacements A grid of 4 rows by 3 columns. 3.10 constructs and their 3.11 replacements 3.10 construct 3.11 replacement behaviour change gather(*coros) TaskGroup siblings cancelled; ExceptionGroup gather(return_exceptions=True) keep, or catch in children none if kept wait_for / async_timeout asyncio.timeout() one deadline for a block except SomeError around gather except* SomeError matches inside groups Each row is a behaviour change at the call site; review them as such, not as renames.

5. Roll out in reviewable steps

Doing all of this in one change makes failures hard to attribute. A sequence that keeps each step small:

  1. Raise the floor in packaging metadata and CI, with no code changes. Delete version branches for 3.10.
  2. Swap timeouts. wait_for and async-timeout → asyncio.timeout(). Lowest risk, easy to test.
  3. Convert one module's gathers to TaskGroup with matching except* handlers, plus tests for the failure paths.
  4. Repeat per module, watching error-rate dashboards between steps: a spike in unhandled ExceptionGroup means a handler was missed.
  5. Add structured-concurrency lint rules — forbid new bare gather without return_exceptions in reviewed code, or flag it for a comment explaining why.
# a small guard for step 4: surface unhandled groups clearly in logs
def log_exception_group(eg: BaseExceptionGroup) -> None:
    for exc in eg.exceptions:
        if isinstance(exc, BaseExceptionGroup):
            log_exception_group(exc)
        else:
            log.error("task failed inside group", exc_info=(type(exc), exc, exc.__traceback__))

Logging nested groups with every child traceback is covered in logging ExceptionGroups with full tracebacks.

Verify: after each step, the error rate and the count of unhandled ExceptionGroup exceptions are unchanged.

A rollout that keeps each change reviewable A flow of 4 stages. A rollout that keeps each change reviewable raise the floor no code changes swap timeouts wait_for to timeout() one module's gathers with except* handlers repeat, then lint watch unhandled groups Small steps make it obvious which change introduced an unhandled ExceptionGroup.

Verification

The adoption is complete when:

  • No gather without return_exceptions remains where all tasks must succeed.
  • Every TaskGroup call site's error handling uses except* or catches inside children.
  • async-timeout is gone and every deadline uses asyncio.timeout().
  • Failure-path tests assert the new exception types and that siblings are cancelled.

Diagnostic Hook: count unhandled ExceptionGroup and BaseExceptionGroup exceptions reaching your top-level error handler, by endpoint. During the rollout that count is the regression signal — a non-zero value means a converted call site whose caller still expects a plain exception. After the rollout, it should stay at zero.

Pitfalls & edge cases

  • Catching the old exception type around a TaskGroup. except NotFound never matches an ExceptionGroup.
  • Assuming TaskGroup returns results. Keep the task objects and call .result() after the block.
  • Converting best-effort fan-outs. If siblings should finish despite a failure, a TaskGroup is the wrong tool.
  • return, break or continue inside except*. All three are a SyntaxError there; set a flag in the handler and act on it after the try.

Frequently Asked Questions

What is the difference between asyncio.gather and TaskGroup?

When one task fails, gather raises the first exception but leaves the other tasks running unobserved. A TaskGroup cancels the remaining tasks, waits for them, and raises an ExceptionGroup containing every failure. In testing, gather's siblings finished after the error; TaskGroup's were cancelled.

Can I use TaskGroup on Python 3.10?

Only through a backport such as the taskgroup package, and except* syntax is still unavailable on 3.10. Raising the minimum version to 3.11 is usually simpler than maintaining both.

Should I replace asyncio.wait_for with asyncio.timeout?

Yes, once the minimum is 3.11. asyncio.timeout bounds a whole block with one deadline, can be rescheduled, and behaves the same on every version, while wait_for ran in a separate task before 3.12.

How do I keep gather's return_exceptions behaviour with TaskGroup?

Either keep gather with return_exceptions=True, which is still correct, or catch expected per-item errors inside each child so the TaskGroup only sees failures that should abort the group.