Leaks from Exceptions Holding Frames¶
An exception object is small; what it references is not. Its __traceback__ links to every frame the exception passed through, and each frame keeps all of its local variables alive — request bodies, query results, buffers. Keep the exception and you keep all of that. Measured on Python 3.14 with a handler that had a 1 MB buffer as a local when it raised: storing the last 50 exceptions for a status page held 47.7 MiB; storing their formatted tracebacks as strings held 0.0 MiB; storing the exceptions after traceback.clear_frames() or with_traceback(None) also held 0.0 MiB. The same 47.7 MiB was held by the results list of asyncio.gather(..., return_exceptions=True) for as long as the list was alive. A long-running task that kept its most recent error in a local variable held one failed request's locals, 1.0 MiB, until it moved on. This guide keeps the information from exceptions without keeping their frames.
Prerequisites¶
- Python 3.11+, stdlib only.
- Memory profiling, from finding memory leaks in asyncio with tracemalloc.
- Reference cycles, from finding reference cycles that keep tasks alive.
1. See what a stored exception keeps alive¶
Storing exceptions is natural — "last N errors" on a debug page, an error list in a batch result, a retry helper that remembers attempts:
recent_errors: collections.deque = collections.deque(maxlen=50)
async def handle(request):
body = await request.read() # possibly megabytes
try:
return await process(body)
except Exception as exc:
recent_errors.append(exc) # keeps exc -> traceback -> this frame -> body
raise
Measured: 50 stored exceptions from a handler with a 1 MB local held 47.7 MiB; after clearing the list, 0.0 MiB. The maxlen bounds the count but not the size, because each exception's cost is whatever its frames' locals happen to be — sometimes nothing, sometimes a 50 MB upload. The chain also continues through __context__ and __cause__ to earlier exceptions and their frames.
Verify: take a tracemalloc snapshot, trigger a batch of failures, take another; growth attributed to request-handling code that persists after the requests completed points at stored exceptions.
2. Store formatted text instead of exception objects¶
For status pages, error reports and batch summaries, a string is what is needed anyway:
import traceback
from dataclasses import dataclass
@dataclass(frozen=True)
class ErrorRecord:
when: float
kind: str
message: str
traceback: str
def record(exc: BaseException) -> ErrorRecord:
return ErrorRecord(
when=time.time(),
kind=type(exc).__qualname__,
message=str(exc),
traceback="".join(traceback.format_exception(exc)), # text: no frame references
)
recent_errors.append(record(exc))
Measured: 50 formatted tracebacks held 0.0 MiB beyond the strings themselves. traceback.TracebackException.from_exception(exc) is a middle ground — a structured, picklable summary without frame references — useful when you want to render the traceback later in different formats. Loggers already format at emit time, so log.exception(...) does not retain anything once the record is handled.
Verify: the status page shows the same information, and memory after a burst of errors returns to baseline.
3. Strip frames when the object must be kept¶
Sometimes the exception object itself must survive: re-raising it later, returning it to a caller, matching on its type. Remove its frames first:
import traceback
# Option A: keep the traceback chain but clear frame locals
traceback.clear_frames(exc.__traceback__)
# Option B: drop the traceback entirely
exc = exc.with_traceback(None)
exc.__context__ = None # earlier exceptions in the chain hold frames too
exc.__cause__ = None
Both measured 0.0 MiB for 50 kept exceptions. clear_frames keeps the traceback's line information for later formatting while releasing every frame's local variables (it cannot clear frames that are still executing); with_traceback(None) keeps only the type and message. Do this at the point where an exception becomes long-lived — when it is put into a cache, a results list or an instance attribute — not in every handler.
Verify: exceptions stored long-term have __traceback__ set to None or point to cleared frames, checked in a unit test.
4. Release gather results and per-task exceptions¶
asyncio.gather(..., return_exceptions=True) puts exceptions into the results list, and each of them carries its task's frames. Measured: the results of 50 failed calls held 47.7 MiB for as long as the list existed. Summarize and drop them:
async def fetch_all(urls: list[str]) -> tuple[list[bytes], list[str]]:
results = await asyncio.gather(*(fetch(u) for u in urls), return_exceptions=True)
ok = [r for r in results if not isinstance(r, BaseException)]
errors = [f"{u}: {r!r}" for u, r in zip(urls, results) if isinstance(r, BaseException)]
return ok, errors # `results` and its exceptions go out of scope here
Returning summaries instead of the raw list lets the exceptions — and their frames — be freed as soon as the function returns; in the test, memory returned to zero after the list was released and the loop ran once. The same applies to TaskGroup failures caught as ExceptionGroup and kept: keep the formatted group, not the group, as in logging ExceptionGroups with full tracebacks.
Verify: functions that use return_exceptions=True do not return or store the raw results list.
5. Watch long-lived coroutines' local variables¶
A coroutine that runs for the life of the service keeps every local variable alive between iterations, including the last exception it saw:
async def poller() -> None:
last_error: str | None = None # a string, not the exception
while True:
try:
await poll_once()
last_error = None
except Exception as exc:
last_error = repr(exc) # measured: keeping `exc` held 1.0 MiB of locals
log.warning("poll failed: %s", last_error)
await asyncio.sleep(INTERVAL)
Measured: a running task that held its last exception in a local kept the failed call's 1 MB of locals alive until the next iteration replaced it. Python deletes the name bound by except ... as exc at the end of the except block precisely to avoid this, but assigning it to another variable bypasses that. In long-running loops, keep a description of the error, not the error.
Verify: a heap snapshot of an idle service contains no exception objects held by running coroutines.
Verification¶
Exceptions do not hold memory when:
- Long-lived error records store text, not exception objects.
- Kept exception objects have frames cleared and their chain cut.
return_exceptions=Trueresults are summarized and released.- Long-running coroutines keep
repr(exc), notexc.
Diagnostic Hook: count live exception objects in a heap snapshot (sum(isinstance(o, BaseException) for o in gc.get_objects())) on an idle instance. A number that grows with the error rate means exceptions are being kept; follow gc.get_referrers from one of them to the list, deque or attribute that holds it.
Pitfalls & edge cases¶
- "Last N errors" as exception objects. Measured: 47.7 MiB for 50.
- Keeping
gatherresults with exceptions. Same 47.7 MiB while the list lives. - Clearing the traceback but not the chain.
__context__and__cause__hold frames too. - Assigning
excto a longer-lived variable. It defeats Python's automatic deletion.
Frequently Asked Questions¶
Why does storing exceptions use so much memory in Python?
Each exception's traceback references every frame it passed through, and each frame keeps its local variables alive. In testing, 50 stored exceptions from a handler with a 1 MB local held 47.7 MiB.
How do I keep an exception without keeping its frames?
Call traceback.clear_frames(exc.traceback) or use exc.with_traceback(None), and clear context and cause; both kept 0 MiB in testing. Or store the formatted traceback as text.
Does asyncio.gather with return_exceptions=True leak memory?
The results list holds exceptions with their frames for as long as it is alive; in testing 50 failures held 47.7 MiB. Summarize the results and release the list.
Does logging an exception keep it in memory?
No. log.exception formats the traceback when the record is handled; nothing is retained afterwards unless a handler stores the record object itself.
Related¶
- Memory & Resource Leaks — up to the topic overview.
- Leaks from forgotten callbacks and timer handles — another way closures keep data alive.
- Resilience, Cancellation & Error Handling — the section overview.