Skip to content

Comparing gather and TaskGroup Cancellation

asyncio.gather and asyncio.TaskGroup both run coroutines concurrently, but they disagree about what happens to the other children when something goes wrong — and the differences decide whether work keeps running after the caller has given up. Measured on Python 3.14 with three children of 1 second each: when one child raised at 0.1 s, gather raised the error to its caller at 0.1 s while the two siblings kept running and finished at 1.0 s, their results discarded. TaskGroup cancelled the siblings, waited for their 0.2 s cleanup, and raised an ExceptionGroup at 0.3 s, with nothing left running. When one child was cancelled directly, gather raised CancelledError in the caller at 0.1 s — as if the caller had been cancelled — while the others ran on; TaskGroup treated it as that child's outcome and returned the other two results at 1.0 s. When the caller itself was cancelled, both cancelled every child and waited for their 0.3 s cleanup, raising at 0.5 s. This guide lays out the cases and when each behaviour is wanted.

Prerequisites

1. Measure what happens when one child fails

Three children: A and C sleep for 1 second, B raises ValueError after 0.1 s. Each child logs when it finishes or is cancelled:

async def worker(name, duration, fail=False, cleanup=0.0):
    try:
        await asyncio.sleep(duration)
        if fail:
            raise ValueError(f"{name} failed")
        log(f"{name} finished")
        return name
    except asyncio.CancelledError:
        await asyncio.sleep(cleanup)              # e.g. rolling back, closing a connection
        log(f"{name} cancelled")
        raise

await asyncio.gather(worker("A", 1.0), worker("B", 0.1, fail=True), worker("C", 1.0))

Measured: gather raised ValueError to the caller at 0.1 s. A and C were not cancelled: both logged "finished" at 1.0 s, 0.9 s after the caller had moved on, and their return values went nowhere. If A and C wrote to a database or called an API, those effects happened after the request that started them had already failed. Nothing in the caller still references them, so their outcome is never observed.

Verify: for each gather call, decide whether siblings should outlive a failure; if not, it should be a TaskGroup.

Three 1-second children, Python 3.14 A grid of 3 rows by 3 columns. Three 1-second children, Python 3.14 event gather TaskGroup child B raises at 0.1 s ValueError at 0.1 s; A, C finish at 1.0 s siblings cancelled; ExceptionGroup at 0.3 s child B cancelled at 0.1 s CancelledError in caller at 0.1 s; A, C finish at 1.0 s returns [A, cancelled, C] at 1.0 s caller cancelled at 0.2 s all cancelled; CancelledError at 0.5 s all cancelled; CancelledError at 0.5 s Cleanup in cancelled children took 0.2-0.3 s.

2. See TaskGroup cancel and wait

The same children in a TaskGroup, with each cancelled child spending 0.2 s on cleanup:

async with asyncio.TaskGroup() as tg:
    tg.create_task(worker("A", 1.0, cleanup=0.2))
    tg.create_task(worker("B", 0.1, fail=True))
    tg.create_task(worker("C", 1.0, cleanup=0.2))

Measured: B's failure at 0.1 s cancelled A and C, both ran their cleanup and logged "cancelled" at 0.3 s, and only then did the async with block exit with ExceptionGroup([ValueError('B failed')]). Nothing ran after the caller received the error. The cost is that the caller waits for cleanup — here 0.2 s — which must itself be bounded, as in bounding cleanup time during cancellation. The error arrives wrapped in an ExceptionGroup, so callers catch it with except*.

Verify: after a TaskGroup raises, no task it created is still running.

3. Watch a directly cancelled child

A child can be cancelled by something other than its group — a per-item timeout, a supervisor, a user action:

tasks = [asyncio.ensure_future(worker(n, 1.0)) for n in "ABC"]
loop.call_later(0.1, tasks[1].cancel)
results = await asyncio.gather(*tasks)

Measured: gather raised CancelledError in the caller at 0.1 s, and A and C finished at 1.0 s in the background. The caller cannot tell this apart from being cancelled itself, so code that treats CancelledError as "I am being shut down" will shut down. With return_exceptions=True, the cancelled child appears as a CancelledError object in the results instead. A TaskGroup treated B's cancellation as B's outcome and nothing more: the group waited for A and C and completed normally at 1.0 s, with B reporting cancelled().

async with asyncio.TaskGroup() as tg:
    tasks = [tg.create_task(worker(n, 1.0)) for n in "ABC"]
    loop.call_later(0.1, tasks[1].cancel)
results = [t.result() if not t.cancelled() else None for t in tasks]

Verify: code that cancels individual children runs them in a TaskGroup, or uses return_exceptions=True and checks each result.

One child fails at 0.1 s (units of 0.1 s) 4 lanes over time. One child fails at 0.1 s (units of 0.1 s) gather: caller waiting moved on at 0.1 s gather: A and C running to 1.0 s, results discarded TaskGroup: caller waiting ExceptionGroup at 0.3 s TaskGroup: A and C running cleanup time → gather's caller leaves early; TaskGroup's children leave with it.

4. Cancel the caller and compare

When the task awaiting the group is cancelled — a client disconnect, a shutdown — both constructs cancel their children. Measured with the caller cancelled at 0.2 s and children spending 0.3 s on cleanup: gather and TaskGroup alike cancelled all three, waited for their cleanup, and raised CancelledError in the caller at 0.5 s. This is the one case where they agree:

async def handle_request():
    async with asyncio.TaskGroup() as tg:          # or gather(); same outcome here
        tg.create_task(fetch_profile(user))
        tg.create_task(fetch_orders(user))

task = asyncio.create_task(handle_request())
await asyncio.sleep(0.2)
task.cancel()                                     # both cancel children and wait for cleanup

The waiting is a feature — cleanup runs before the caller continues — but it means cancellation is not instant. A child whose cleanup hangs holds the caller's cancellation indefinitely, with either construct.

Verify: cancelling a request handler that fans out leaves no children running, and the handler finishes within the cleanup bound.

5. Choose deliberately

The measurements reduce to a rule. Use a TaskGroup when the children are parts of one operation — if one fails, the rest are wasted work or must not happen. Use gather(return_exceptions=True) when the children are independent and you want every result, success or failure:

# Parts of one answer: all or nothing
async with asyncio.TaskGroup() as tg:
    profile = tg.create_task(fetch_profile(user_id))
    orders = tg.create_task(fetch_orders(user_id))

# Independent deliveries: report each outcome
results = await asyncio.gather(*(notify(u) for u in users), return_exceptions=True)
failed = [u for u, r in zip(users, results) if isinstance(r, BaseException)]

Plain gather without return_exceptions is rarely the right choice: it reports the first error and leaves the rest running unobserved. When independent tasks should not cancel each other but still need structured lifetime, collect errors inside each child instead, as in collecting errors without cancelling siblings. For moving existing code, see migrating from gather to TaskGroup.

gather or TaskGroup? A decision on How are the children related with 3 outcomes. gather or TaskGroup? How are the children related? parts of one operation TaskGroup failure cancels siblings independent, want every outcome gather(return_exceptions=True) check each result independent, structured lifetime TaskGroup + catch inside each child no sibling cancels another Plain gather without return_exceptions fits none of these.

Verify: every remaining plain gather call has a comment explaining why siblings may outlive a failure.

Verification

Concurrency constructs are chosen correctly when:

  • Parts of one operation run in a TaskGroup, so a failure cancels the rest.
  • Independent work uses gather(return_exceptions=True) and checks each result.
  • Individually cancelled children do not masquerade as the caller's own cancellation.
  • Cleanup in cancelled children is bounded, since both constructs wait for it.

Diagnostic Hook: when database writes or API calls appear after a request has already failed, look for a plain gather in that request. With one child failing at 0.1 s, its siblings ran on and finished at 1.0 s in this test, after the caller had returned the error.

Pitfalls & edge cases

  • Plain gather for dependent work. Measured: siblings ran 0.9 s after the caller left.
  • Cancelling one child of a gather. Measured: the caller got CancelledError.
  • Unbounded cleanup. TaskGroup and gather both wait for it.
  • Catching ValueError around a TaskGroup. It arrives inside an ExceptionGroup.

Frequently Asked Questions

Does asyncio.gather cancel other tasks when one fails?

No. The first exception is raised to the caller at once and the other tasks keep running: siblings finished at 1.0 s after gather raised at 0.1 s. TaskGroup cancels them.

Why does gather raise CancelledError when I cancel one task?

gather propagates a child's cancellation to its caller. Cancelling one child at 0.1 s raised CancelledError in the caller while the others ran on.

Is TaskGroup slower to report errors than gather?

It waits for cancelled siblings to clean up: the error arrived at 0.3 s instead of 0.1 s with 0.2 s of cleanup. In exchange, nothing keeps running.

When should I still use asyncio.gather?

For independent tasks with return_exceptions=True, when you want every outcome and one failure should not stop the others.