Skip to content

Leaks from Exceptions Holding Frames

An exception object is small; what it references is not. Its __traceback__ links to every frame the exception passed through, and each frame keeps all of its local variables alive — request bodies, query results, buffers. Keep the exception and you keep all of that. Measured on Python 3.14 with a handler that had a 1 MB buffer as a local when it raised: storing the last 50 exceptions for a status page held 47.7 MiB; storing their formatted tracebacks as strings held 0.0 MiB; storing the exceptions after traceback.clear_frames() or with_traceback(None) also held 0.0 MiB. The same 47.7 MiB was held by the results list of asyncio.gather(..., return_exceptions=True) for as long as the list was alive. A long-running task that kept its most recent error in a local variable held one failed request's locals, 1.0 MiB, until it moved on. This guide keeps the information from exceptions without keeping their frames.

Prerequisites

1. See what a stored exception keeps alive

Storing exceptions is natural — "last N errors" on a debug page, an error list in a batch result, a retry helper that remembers attempts:

recent_errors: collections.deque = collections.deque(maxlen=50)


async def handle(request):
    body = await request.read()            # possibly megabytes
    try:
        return await process(body)
    except Exception as exc:
        recent_errors.append(exc)          # keeps exc -> traceback -> this frame -> body
        raise

Measured: 50 stored exceptions from a handler with a 1 MB local held 47.7 MiB; after clearing the list, 0.0 MiB. The maxlen bounds the count but not the size, because each exception's cost is whatever its frames' locals happen to be — sometimes nothing, sometimes a 50 MB upload. The chain also continues through __context__ and __cause__ to earlier exceptions and their frames.

Verify: take a tracemalloc snapshot, trigger a batch of failures, take another; growth attributed to request-handling code that persists after the requests completed points at stored exceptions.

Memory held by keeping 50 errors from a handler with a 1 MB local 5 horizontal bars comparing exception objects with the others. Memory held by keeping 50 errors from a handler with a 1 MB local exception objects 47.7 MiB gather(return_exceptions=True) results 47.7 MiB formatted traceback strings 0.0 MiB after traceback.clear_frames() 0.0 MiB with_traceback(None) 0.0 MiB Python 3.14; tracemalloc totals after 50 failures. Keep the text of a traceback, not its frames.

2. Store formatted text instead of exception objects

For status pages, error reports and batch summaries, a string is what is needed anyway:

import traceback
from dataclasses import dataclass


@dataclass(frozen=True)
class ErrorRecord:
    when: float
    kind: str
    message: str
    traceback: str


def record(exc: BaseException) -> ErrorRecord:
    return ErrorRecord(
        when=time.time(),
        kind=type(exc).__qualname__,
        message=str(exc),
        traceback="".join(traceback.format_exception(exc)),   # text: no frame references
    )


recent_errors.append(record(exc))

Measured: 50 formatted tracebacks held 0.0 MiB beyond the strings themselves. traceback.TracebackException.from_exception(exc) is a middle ground — a structured, picklable summary without frame references — useful when you want to render the traceback later in different formats. Loggers already format at emit time, so log.exception(...) does not retain anything once the record is handled.

Verify: the status page shows the same information, and memory after a burst of errors returns to baseline.

3. Strip frames when the object must be kept

Sometimes the exception object itself must survive: re-raising it later, returning it to a caller, matching on its type. Remove its frames first:

import traceback

# Option A: keep the traceback chain but clear frame locals
traceback.clear_frames(exc.__traceback__)

# Option B: drop the traceback entirely
exc = exc.with_traceback(None)
exc.__context__ = None                     # earlier exceptions in the chain hold frames too
exc.__cause__ = None

Both measured 0.0 MiB for 50 kept exceptions. clear_frames keeps the traceback's line information for later formatting while releasing every frame's local variables (it cannot clear frames that are still executing); with_traceback(None) keeps only the type and message. Do this at the point where an exception becomes long-lived — when it is put into a cache, a results list or an instance attribute — not in every handler.

Verify: exceptions stored long-term have __traceback__ set to None or point to cleared frames, checked in a unit test.

What one stored exception can keep alive A flow of 5 stages. What one stored exception can keep alive stored exception in a list or attribute __traceback__ every frame passed frame locals bodies, results, buffers __context__ / __cause__ earlier exceptions their frames more locals Measured: 1 MB of locals per exception, multiplied by every one kept.

4. Release gather results and per-task exceptions

asyncio.gather(..., return_exceptions=True) puts exceptions into the results list, and each of them carries its task's frames. Measured: the results of 50 failed calls held 47.7 MiB for as long as the list existed. Summarize and drop them:

async def fetch_all(urls: list[str]) -> tuple[list[bytes], list[str]]:
    results = await asyncio.gather(*(fetch(u) for u in urls), return_exceptions=True)
    ok = [r for r in results if not isinstance(r, BaseException)]
    errors = [f"{u}: {r!r}" for u, r in zip(urls, results) if isinstance(r, BaseException)]
    return ok, errors                      # `results` and its exceptions go out of scope here

Returning summaries instead of the raw list lets the exceptions — and their frames — be freed as soon as the function returns; in the test, memory returned to zero after the list was released and the loop ran once. The same applies to TaskGroup failures caught as ExceptionGroup and kept: keep the formatted group, not the group, as in logging ExceptionGroups with full tracebacks.

Verify: functions that use return_exceptions=True do not return or store the raw results list.

5. Watch long-lived coroutines' local variables

A coroutine that runs for the life of the service keeps every local variable alive between iterations, including the last exception it saw:

async def poller() -> None:
    last_error: str | None = None             # a string, not the exception
    while True:
        try:
            await poll_once()
            last_error = None
        except Exception as exc:
            last_error = repr(exc)            # measured: keeping `exc` held 1.0 MiB of locals
            log.warning("poll failed: %s", last_error)
        await asyncio.sleep(INTERVAL)

Measured: a running task that held its last exception in a local kept the failed call's 1 MB of locals alive until the next iteration replaced it. Python deletes the name bound by except ... as exc at the end of the except block precisely to avoid this, but assigning it to another variable bypasses that. In long-running loops, keep a description of the error, not the error.

Verify: a heap snapshot of an idle service contains no exception objects held by running coroutines.

How should this code keep an error? A decision on Why is the error being kept with 4 outcomes. How should this code keep an error? Why is the error being kept? show or report it format to text 0 MiB retained re-raise or match later clear_frames / with_traceback(None) keep type + message gather return_exceptions summarize, drop the list frames freed long-running loop keep repr(exc) not the object An exception's cost is its frames' locals; keep only what you need.

Verification

Exceptions do not hold memory when:

  • Long-lived error records store text, not exception objects.
  • Kept exception objects have frames cleared and their chain cut.
  • return_exceptions=True results are summarized and released.
  • Long-running coroutines keep repr(exc), not exc.

Diagnostic Hook: count live exception objects in a heap snapshot (sum(isinstance(o, BaseException) for o in gc.get_objects())) on an idle instance. A number that grows with the error rate means exceptions are being kept; follow gc.get_referrers from one of them to the list, deque or attribute that holds it.

Pitfalls & edge cases

  • "Last N errors" as exception objects. Measured: 47.7 MiB for 50.
  • Keeping gather results with exceptions. Same 47.7 MiB while the list lives.
  • Clearing the traceback but not the chain. __context__ and __cause__ hold frames too.
  • Assigning exc to a longer-lived variable. It defeats Python's automatic deletion.

Frequently Asked Questions

Why does storing exceptions use so much memory in Python?

Each exception's traceback references every frame it passed through, and each frame keeps its local variables alive. In testing, 50 stored exceptions from a handler with a 1 MB local held 47.7 MiB.

How do I keep an exception without keeping its frames?

Call traceback.clear_frames(exc.traceback) or use exc.with_traceback(None), and clear context and cause; both kept 0 MiB in testing. Or store the formatted traceback as text.

Does asyncio.gather with return_exceptions=True leak memory?

The results list holds exceptions with their frames for as long as it is alive; in testing 50 failures held 47.7 MiB. Summarize the results and release the list.

Does logging an exception keep it in memory?

No. log.exception formats the traceback when the record is handled; nothing is retained afterwards unless a handler stores the record object itself.