Finding Reference Cycles That Keep Tasks Alive¶
An object that stores its own task — self.task = asyncio.create_task(self.run()) — forms a reference cycle the moment that task fails: the task holds the exception, the exception's traceback holds the coroutine's frame, the frame holds self, and self holds the task. Reference counting cannot free a cycle; only the cyclic garbage collector can, and it runs on its own schedule. Measured on Python 3.14 with the collector paused, as it effectively is between collections: 50 failed jobs, each with a 1 MB buffer and dropped by the code that created them, were all still alive, holding 47.7 MiB; after gc.collect(), 1.0 MiB. Clearing self.task in a done callback broke the cycle and left 0 jobs alive without any collection. The opposite failure exists too: pending tasks referenced only through their own objects were destroyed by the collector, with asyncio's "Task was destroyed but it is pending!" warning. This guide finds both kinds of cycle and removes them.
Prerequisites¶
- Python 3.11+, stdlib only.
- Memory profiling basics, from finding memory leaks in asyncio with tracemalloc.
- Task tracking, from tracking task growth in long-running services.
1. Recognise the task-exception cycle¶
The cycle needs only an object that keeps its task and a task that fails:
class Job:
def __init__(self, payload: bytes) -> None:
self.payload = payload # something large
self.task = asyncio.create_task(self.run())
async def run(self) -> None:
data = self.payload
raise ValueError("job failed")
# job -> task -> exception -> traceback -> frame of run() -> self (job)
Measured with automatic collection paused: 50 such jobs, dropped after failing, all stayed alive with 47.7 MiB traced; gc.collect() freed them down to 1.0 MiB. In a real service the collector runs eventually, so this is a delay rather than a permanent leak — but its timing depends on allocation counts, not memory size, so a few large objects can be held for a long time while memory climbs. Services that disable automatic collection or tune its thresholds for latency make it worse.
Verify: gc.collect() in a debug endpoint frees a large amount of memory; if so, cycles are holding it between collections.
2. Break the cycle when the task finishes¶
Clear the stored reference when the task is done, and the cycle never forms in a way that survives:
class Job:
def __init__(self, payload: bytes) -> None:
self.payload = payload
self.task: asyncio.Task | None = asyncio.create_task(self.run())
self.task.add_done_callback(self._on_done)
def _on_done(self, task: asyncio.Task) -> None:
if not task.cancelled() and task.exception() is not None:
log.error("job failed", exc_info=task.exception())
self.task = None # break job -> task
Measured: with the done callback, all 50 failed jobs were freed immediately — 0 alive without any collection. Retrieving task.exception() in the callback also marks the exception as retrieved, so asyncio does not log "Task exception was never retrieved" later. Two alternatives work the same way: keep tasks in an external registry (a set with discard as the done callback) instead of on the object, or catch exceptions inside run() so the task never stores one.
Verify: after a burst of failing jobs, len(gc.get_objects()) of the job type returns to baseline without a manual collection.
3. Find which objects are in cycles¶
When memory is held until a collection, find out what is involved. The collector can save what it would have freed for inspection:
import gc
from collections import Counter
def cycle_report(top: int = 10) -> list[tuple[str, int]]:
gc.collect()
gc.set_debug(gc.DEBUG_SAVEALL) # keep unreachable objects in gc.garbage
try:
gc.collect()
counts = Counter(type(o).__qualname__ for o in gc.garbage)
return counts.most_common(top)
finally:
gc.set_debug(0)
gc.garbage.clear()
# e.g. [('frame', 412), ('Task', 50), ('ValueError', 50), ('traceback', 150), ('Job', 50)]
Run it once after a period of normal traffic, from a debug endpoint or a test, not continuously — DEBUG_SAVEALL keeps everything it finds. Types that show up together in the report (your class, Task, an exception type, traceback, frame) are the cycle's members. To see the exact chain for one object, gc.get_referrers() walks backwards, or the objgraph package draws it.
Verify: the report after your fix no longer lists your class or Task.
4. Watch for the opposite: tasks that vanish¶
asyncio's event loop keeps only weak references to tasks. A task referenced solely through a cycle — its object holds it, it holds its object, nothing else holds either — is garbage the collector may destroy while still pending:
class Poller:
def __init__(self) -> None:
self.task = asyncio.create_task(self.loop()) # the only strong reference
async def loop(self) -> None:
while True:
await poll_once()
Poller() # nobody keeps it: measured, the collector destroyed it while pending
Tested: five such objects were collected, and asyncio logged "Task was destroyed but it is pending!" — background work silently stopped. This is the same root cause as fire-and-forget tasks disappearing, and the fix is the same: keep a strong reference somewhere with a lifetime you control, such as a set of background tasks owned by the application, as in naming and tracking tasks for observability, or a TaskGroup that owns them.
Verify: grep logs for "Task was destroyed but it is pending"; any occurrence is a task that lost its last strong reference.
5. Guard against regressions in tests¶
Cycles creep back in with refactoring. A test can check that an object is freed by reference counting alone:
import gc
import weakref
async def test_failed_job_is_freed_without_gc():
gc.disable()
try:
job = Job(b"x" * 1_000_000)
with pytest.raises(ValueError):
await job.task
ref = weakref.ref(job)
del job
await asyncio.sleep(0) # let done callbacks run
assert ref() is None, "Job is kept alive by a reference cycle"
finally:
gc.enable()
Disabling the collector for the test makes cycles visible as failures instead of passing by luck when a collection happens to run. Pair it with a check that no "Task was destroyed but it is pending" warning was logged, so both directions — objects held too long and tasks freed too soon — stay covered.
Verify: the test fails when the done callback that clears self.task is removed.
Verification¶
Task references are healthy when:
- Objects that store their tasks clear them when done.
- Background tasks have an application-owned strong reference.
- A cycle report shows no task or job types after normal traffic.
- Tests with the collector disabled confirm objects are freed by reference counting.
Diagnostic Hook: export the amount freed by periodic gc.collect() calls (gc.callbacks can measure collected counts) and the count of "Task was destroyed but it is pending" warnings. Large amounts freed per collection mean cycles are holding memory between collections; any destroyed-pending warning means a task lost its last owner.
Pitfalls & edge cases¶
self.taskon objects whose tasks can fail. Measured: 47.7 MiB held until collection.- Relying on the collector's timing. It counts allocations, not bytes.
- Tasks owned only through cycles. Tested: destroyed while pending.
- Leaving
DEBUG_SAVEALLon. It keeps everything it finds.
Frequently Asked Questions¶
Why does memory drop when I call gc.collect() in my asyncio service?
Objects are held in reference cycles, commonly an object storing its own task whose exception's traceback references the object. In testing, 47.7 MiB from 50 failed jobs was freed only by gc.collect().
How do I avoid reference cycles with asyncio tasks?
Clear the stored task reference in a done callback, keep tasks in an external registry instead of on the object, or handle exceptions inside the coroutine so the task does not store one.
What does 'Task was destroyed but it is pending' mean?
The task lost its last strong reference and was garbage-collected before finishing; the event loop only holds weak references. Keep tasks in an application-owned set or TaskGroup.
How do I find reference cycles in Python?
Set gc.DEBUG_SAVEALL, run gc.collect(), and count the types in gc.garbage; types that appear together form the cycle. Clear gc.garbage and the debug flag afterwards.
Related¶
- Memory & Resource Leaks — up to the topic overview.
- Leaks from exceptions holding frames — the exception half of this cycle in detail.
- Resilience, Cancellation & Error Handling — the section overview.