Collecting Errors Without Cancelling Siblings¶
A TaskGroup is all-or-nothing: the first exception cancels every sibling. For independent jobs — notifications, per-tenant reports, health probes — that is wrong, and the goal is the opposite: let every job finish and report every failure together. Measured on Python 3.14 with ten jobs, three of which raised different exceptions at 0.1, 0.3 and 0.5 s: a plain TaskGroup raised at 0.10 s with one error and 0 of the ten jobs completed — the other two failures never happened, because their jobs were cancelled first. Catching inside each child and raising one ExceptionGroup at the end took 0.65 s, completed 7 jobs and reported all 3 errors — ValueError, KeyError and ConnectionError — each with a note naming its job. gather(return_exceptions=True) gave the same counts. Cancelling the collecting version from outside still propagated CancelledError. This guide builds the collecting pattern and shows how callers handle its result.
Prerequisites¶
- Python 3.11+ for
TaskGroup,ExceptionGroupandadd_note. - Group handling, from handling specific errors with except*.
- The topic overview, Exception Groups & TaskGroups.
1. See what a plain TaskGroup reports¶
Ten independent jobs, three of which fail at different times:
async def job(i):
if i in FAIL:
delay, exc = FAIL[i] # 2: 0.1 s ValueError, 5: 0.3 s KeyError, 8: 0.5 s ConnectionError
await asyncio.sleep(delay)
raise exc(f"job {i}")
await asyncio.sleep(0.2 + i * 0.05)
return i
async with asyncio.TaskGroup() as tg:
for i in range(10):
tg.create_task(job(i))
Measured: the group raised an ExceptionGroup containing one ValueError at 0.10 s. No job had completed yet — the earliest successful one would have finished at 0.2 s — so all seven healthy jobs were cancelled, along with the two that would have failed later. The caller learned about one failure out of three, and nothing else. For one operation split into parts, that is exactly right; for ten independent jobs, it discards nine results to report one error.
Verify: for each TaskGroup, ask whether one child's failure makes the others' results useless; if not, collect errors instead.
2. Catch inside each child and raise one group¶
Keep the TaskGroup for its structure — no child outlives the block, and outside cancellation reaches every child — but stop individual failures from reaching it:
async def run_all(jobs) -> list:
results, errors = [], []
async def guarded(i, coro):
try:
results.append(await coro)
except Exception as e: # Exception only: cancellation must propagate
e.add_note(f"while running job {i}")
errors.append(e)
async with asyncio.TaskGroup() as tg:
for i, coro in enumerate(jobs):
tg.create_task(guarded(i, coro))
if errors:
raise ExceptionGroup(f"{len(errors)} of {len(jobs)} jobs failed", errors)
return results
Measured: all ten jobs ran to their natural end at 0.65 s, seven succeeded, and the ExceptionGroup carried the ValueError, KeyError and ConnectionError, each with its note. The exception objects keep their tracebacks, so a log of the group shows where each failure happened; for large fan-outs, mind the memory that holds, as measured in collecting errors from worker pools.
Verify: a test with several failing jobs receives every failure in one group, and every successful result.
3. Handle the group with except*¶
Callers split the group by exception type, handling each kind once no matter how many jobs raised it:
try:
results = await run_all(jobs)
except* ConnectionError as group:
for e in group.exceptions:
log.warning("transient: %s (%s)", e, e.__notes__)
schedule_retry([job_for(e) for e in group.exceptions])
except* (ValueError, KeyError) as group:
log.error("bad input in %d jobs", len(group.exceptions))
Measured: the ConnectionError handler received exactly one exception, job 8, with the note while running job 8; the ValueError/KeyError handler received the other two. Each except* clause runs at most once, with a sub-group of the matching exceptions; unmatched exceptions continue to propagate in a new group. Results from successful jobs are lost when the group is raised, so if callers need both, return them alongside the errors instead of raising — step 4.
Verify: each except* clause handles its whole sub-group, and an unexpected exception type still propagates.
4. Return results and errors together when both matter¶
Raising discards the successful results. When callers need both — a report that lists what succeeded and what failed — return a structure instead:
@dataclass
class Outcome:
results: dict[int, object]
errors: dict[int, Exception]
async def run_all_outcomes(jobs) -> Outcome:
out = Outcome({}, {})
async def guarded(i, coro):
try:
out.results[i] = await coro
except Exception as e:
out.errors[i] = e
async with asyncio.TaskGroup() as tg:
for i, coro in enumerate(jobs):
tg.create_task(guarded(i, coro))
return out
This is what gather(return_exceptions=True) gives, as a list in submission order: measured, it also completed seven jobs and returned three exceptions at 0.65 s. The TaskGroup version keys outcomes by job, which survives reordering, and keeps one structural advantage: if a programming error escapes the guard — anything not caught by except Exception — the group still cancels the siblings rather than leaving them running unobserved, as plain gather does in comparing gather and TaskGroup cancellation.
Verify: the outcome structure contains exactly one entry per job, in either results or errors.
5. Keep cancellation working¶
The guard catches Exception, not BaseException, so CancelledError — a BaseException since Python 3.8 — passes through. Measured: cancelling the collecting run at 0.15 s raised CancelledError in the caller, as it should. Two related details:
async def guarded(i, coro):
try:
results.append(await coro)
except Exception as e:
errors.append(e)
# no `except BaseException`, no bare `except:` - KeyboardInterrupt, SystemExit
# and CancelledError must reach the TaskGroup
Do not convert collected exceptions into strings early if callers might need except* matching — the type is the dispatch key. And if a job should be retried rather than reported, decide that inside the guard, before it becomes one entry in a group of hundreds; see making retry loops cancellation-safe.
Verify: cancelling the caller during a collecting run raises CancelledError promptly, with every child cancelled.
Verification¶
Error collection works when:
- Independent jobs all run to completion, regardless of others' failures.
- Every failure is reported, with a note identifying its job.
- Callers can dispatch by type with
except*, or read a results-and-errors structure. - Cancellation still propagates, because only
Exceptionis caught.
Diagnostic Hook: when a batch reports a single error but several items turn out to be broken, check whether the batch runs in a plain TaskGroup. With three failing jobs, the default group reported one and cancelled the rest at 0.10 s in this test.
Pitfalls & edge cases¶
- A plain TaskGroup for independent jobs. Measured: 1 of 3 errors seen, 0 of 10 jobs done.
- Catching
BaseExceptionin the guard. It intercepts cancellation meant for the job. - Raising the group when results are needed. Return an outcome structure instead.
- Stringifying errors early.
except*needs the exception types.
Frequently Asked Questions¶
How do I stop a TaskGroup from cancelling other tasks on error?
Catch Exception inside each child and record it, then raise one ExceptionGroup after the block. Ten jobs ran to completion with all three failures reported.
Is gather(return_exceptions=True) the same as collecting in a TaskGroup?
For outcomes, yes: both completed 7 of 10 jobs and returned 3 errors. The TaskGroup version keys results by job and still cancels siblings on uncaught programming errors.
How do I tell which task an exception came from?
Add a note when catching it: e.add_note(f"while running job {i}"). The note appears in tracebacks and in e.notes.
Does collecting errors break cancellation?
Not if the guard catches Exception only. Cancelling the collecting run at 0.15 s raised CancelledError in the caller.
Related¶
- Exception Groups & TaskGroups — up to the topic overview.
- Adding notes to exceptions in async code — the context each collected error carries.
- Resilience, Cancellation & Error Handling — the section overview.