Skip to content

Collecting Errors Without Cancelling Siblings

A TaskGroup is all-or-nothing: the first exception cancels every sibling. For independent jobs — notifications, per-tenant reports, health probes — that is wrong, and the goal is the opposite: let every job finish and report every failure together. Measured on Python 3.14 with ten jobs, three of which raised different exceptions at 0.1, 0.3 and 0.5 s: a plain TaskGroup raised at 0.10 s with one error and 0 of the ten jobs completed — the other two failures never happened, because their jobs were cancelled first. Catching inside each child and raising one ExceptionGroup at the end took 0.65 s, completed 7 jobs and reported all 3 errors — ValueError, KeyError and ConnectionError — each with a note naming its job. gather(return_exceptions=True) gave the same counts. Cancelling the collecting version from outside still propagated CancelledError. This guide builds the collecting pattern and shows how callers handle its result.

Prerequisites

1. See what a plain TaskGroup reports

Ten independent jobs, three of which fail at different times:

async def job(i):
    if i in FAIL:
        delay, exc = FAIL[i]                   # 2: 0.1 s ValueError, 5: 0.3 s KeyError, 8: 0.5 s ConnectionError
        await asyncio.sleep(delay)
        raise exc(f"job {i}")
    await asyncio.sleep(0.2 + i * 0.05)
    return i

async with asyncio.TaskGroup() as tg:
    for i in range(10):
        tg.create_task(job(i))

Measured: the group raised an ExceptionGroup containing one ValueError at 0.10 s. No job had completed yet — the earliest successful one would have finished at 0.2 s — so all seven healthy jobs were cancelled, along with the two that would have failed later. The caller learned about one failure out of three, and nothing else. For one operation split into parts, that is exactly right; for ten independent jobs, it discards nine results to report one error.

Verify: for each TaskGroup, ask whether one child's failure makes the others' results useless; if not, collect errors instead.

10 independent jobs, 3 failing at 0.1, 0.3 and 0.5 s A grid of 3 rows by 4 columns. 10 independent jobs, 3 failing at 0.1, 0.3 and 0.5 s approach time jobs completed errors reported TaskGroup, default 0.10 s 0 of 10 1 (ValueError) TaskGroup, catch inside each child 0.65 s 7 of 10 3, each with a note gather(return_exceptions=True) 0.65 s 7 of 10 3 Python 3.14.

2. Catch inside each child and raise one group

Keep the TaskGroup for its structure — no child outlives the block, and outside cancellation reaches every child — but stop individual failures from reaching it:

async def run_all(jobs) -> list:
    results, errors = [], []

    async def guarded(i, coro):
        try:
            results.append(await coro)
        except Exception as e:                   # Exception only: cancellation must propagate
            e.add_note(f"while running job {i}")
            errors.append(e)

    async with asyncio.TaskGroup() as tg:
        for i, coro in enumerate(jobs):
            tg.create_task(guarded(i, coro))

    if errors:
        raise ExceptionGroup(f"{len(errors)} of {len(jobs)} jobs failed", errors)
    return results

Measured: all ten jobs ran to their natural end at 0.65 s, seven succeeded, and the ExceptionGroup carried the ValueError, KeyError and ConnectionError, each with its note. The exception objects keep their tracebacks, so a log of the group shows where each failure happened; for large fan-outs, mind the memory that holds, as measured in collecting errors from worker pools.

Verify: a test with several failing jobs receives every failure in one group, and every successful result.

3. Handle the group with except*

Callers split the group by exception type, handling each kind once no matter how many jobs raised it:

try:
    results = await run_all(jobs)
except* ConnectionError as group:
    for e in group.exceptions:
        log.warning("transient: %s (%s)", e, e.__notes__)
    schedule_retry([job_for(e) for e in group.exceptions])
except* (ValueError, KeyError) as group:
    log.error("bad input in %d jobs", len(group.exceptions))

Measured: the ConnectionError handler received exactly one exception, job 8, with the note while running job 8; the ValueError/KeyError handler received the other two. Each except* clause runs at most once, with a sub-group of the matching exceptions; unmatched exceptions continue to propagate in a new group. Results from successful jobs are lost when the group is raised, so if callers need both, return them alongside the errors instead of raising — step 4.

Verify: each except* clause handles its whole sub-group, and an unexpected exception type still propagates.

Collecting errors from independent jobs A flow of 5 stages. Collecting errors from independent jobs Guard each job catch Exception, not BaseException Annotate add_note('while running job 8') Group completes every job ran to its end Report one ExceptionGroup, or results + errors Caller except* by type The TaskGroup still bounds lifetimes and propagates cancellation.

4. Return results and errors together when both matter

Raising discards the successful results. When callers need both — a report that lists what succeeded and what failed — return a structure instead:

@dataclass
class Outcome:
    results: dict[int, object]
    errors: dict[int, Exception]

async def run_all_outcomes(jobs) -> Outcome:
    out = Outcome({}, {})

    async def guarded(i, coro):
        try:
            out.results[i] = await coro
        except Exception as e:
            out.errors[i] = e

    async with asyncio.TaskGroup() as tg:
        for i, coro in enumerate(jobs):
            tg.create_task(guarded(i, coro))
    return out

This is what gather(return_exceptions=True) gives, as a list in submission order: measured, it also completed seven jobs and returned three exceptions at 0.65 s. The TaskGroup version keys outcomes by job, which survives reordering, and keeps one structural advantage: if a programming error escapes the guard — anything not caught by except Exception — the group still cancels the siblings rather than leaving them running unobserved, as plain gather does in comparing gather and TaskGroup cancellation.

Verify: the outcome structure contains exactly one entry per job, in either results or errors.

5. Keep cancellation working

The guard catches Exception, not BaseException, so CancelledError — a BaseException since Python 3.8 — passes through. Measured: cancelling the collecting run at 0.15 s raised CancelledError in the caller, as it should. Two related details:

async def guarded(i, coro):
    try:
        results.append(await coro)
    except Exception as e:
        errors.append(e)
    # no `except BaseException`, no bare `except:` - KeyboardInterrupt, SystemExit
    # and CancelledError must reach the TaskGroup

Do not convert collected exceptions into strings early if callers might need except* matching — the type is the dispatch key. And if a job should be retried rather than reported, decide that inside the guard, before it becomes one entry in a group of hundreds; see making retry loops cancellation-safe.

Verify: cancelling the caller during a collecting run raises CancelledError promptly, with every child cancelled.

Jobs completed out of 10 3 horizontal bars comparing TaskGroup, default (1 error seen) with the others. Jobs completed out of 10 TaskGroup, default (1 error seen) 0 TaskGroup, catch inside (3 errors seen) 7 gather(return_exceptions=True) (3 seen) 7 Collecting turned 1 visible failure into 3, and 0 results into 7.

Verification

Error collection works when:

  • Independent jobs all run to completion, regardless of others' failures.
  • Every failure is reported, with a note identifying its job.
  • Callers can dispatch by type with except*, or read a results-and-errors structure.
  • Cancellation still propagates, because only Exception is caught.

Diagnostic Hook: when a batch reports a single error but several items turn out to be broken, check whether the batch runs in a plain TaskGroup. With three failing jobs, the default group reported one and cancelled the rest at 0.10 s in this test.

Pitfalls & edge cases

  • A plain TaskGroup for independent jobs. Measured: 1 of 3 errors seen, 0 of 10 jobs done.
  • Catching BaseException in the guard. It intercepts cancellation meant for the job.
  • Raising the group when results are needed. Return an outcome structure instead.
  • Stringifying errors early. except* needs the exception types.

Frequently Asked Questions

How do I stop a TaskGroup from cancelling other tasks on error?

Catch Exception inside each child and record it, then raise one ExceptionGroup after the block. Ten jobs ran to completion with all three failures reported.

Is gather(return_exceptions=True) the same as collecting in a TaskGroup?

For outcomes, yes: both completed 7 of 10 jobs and returned 3 errors. The TaskGroup version keys results by job and still cancels siblings on uncaught programming errors.

How do I tell which task an exception came from?

Add a note when catching it: e.add_note(f"while running job {i}"). The note appears in tracebacks and in e.notes.

Does collecting errors break cancellation?

Not if the guard catches Exception only. Cancelling the collecting run at 0.15 s raised CancelledError in the caller.