Tuning Garbage Collection for Event Loop Latency¶
Python's cyclic garbage collector runs on the thread that triggers it — in an asyncio service, the event loop thread — and while it runs, nothing else does. Collections of young objects are short; a full collection scans every tracked object in the process, so its pause grows with the heap. Measured with a service-like workload on CPython 3.13.14 and 3.14.6: a cache of 1,000,000 small dicts held for the life of the process, plus request handling that created reference cycles. Young-generation collections ran tens of thousands of times with a median pause of 0.11 ms; a single full collection during the 15-second run paused the loop for 1,210 ms on 3.14.6 and 1,279 ms on 3.13.14, so the maximum loop lag equalled it. Calling gc.freeze() after building the cache moved those objects out of the collector's reach, and the maximum loop lag fell to 1.1 ms. Raising collection thresholds to (50,000, 50, 100) made young collections bigger — median 3.2 ms, worst 44 ms — and loop lag p99 10.7 ms. A different 3.14 build, which reported thresholds of (2000, 10, 0) and collected incrementally, still paused for up to 844–1,003 ms without freezing and 4.4 ms with it. This guide finds GC pauses and removes the large ones.
Prerequisites¶
- CPython 3.12+; the
gcmodule. - Measuring loop lag, from measuring event loop lag in production.
- The topic overview, Event Loop Configuration.
1. Measure GC pauses directly¶
gc.callbacks lets you time every collection. Record the generation and duration alongside loop lag, so you can tell GC stalls from other blocking:
import gc, time
pauses: list[tuple[int, float]] = []
_started = 0.0
def gc_timer(phase: str, info: dict) -> None:
global _started
if phase == "start":
_started = time.perf_counter()
else:
pauses.append((info["generation"], time.perf_counter() - _started))
gc.callbacks.append(gc_timer)
Measured over 15 s on 3.14.6 with the million-object cache: about 67,000 generation-0 collections with a median of 0.11 ms, about 6,000 generation-1 collections with a median of 0.11 ms, and one generation-2 collection of 1,210 ms. The loop-lag probe's maximum was 1,210 ms — the same event. The young collections added up to several seconds of GC time over the run because the workload created cycles constantly, but no single one mattered for latency; the one full collection did. Export the maximum pause per generation as a metric; averages hide it entirely.
Verify: you have per-generation GC pause maxima from production or a realistic load test, next to loop-lag maxima.
2. Freeze the heap you built at start-up¶
Most of a long-running service's heap is created at start-up and kept: configuration, caches, routing tables, imported modules, ORM metadata. Full collections scan all of it, every time, and find nothing to free. gc.freeze() moves every object currently tracked into a permanent generation that collections skip:
import gc
async def main():
await load_caches() # build the long-lived heap
gc.collect() # free start-up garbage first
gc.freeze() # everything alive now is exempt from future scans
await serve()
Measured on 3.14.6: no full collection ran in the 15-second test, young collections were unchanged, and the maximum loop lag was 1.1 ms instead of 1,210 ms. On the incremental build, freezing cut the worst pause from 844–1,003 ms to 4.4 ms. Frozen objects are still freed by reference counting when their last reference goes; only cycles among them are never collected, so freeze after start-up, not after objects that churn. In pre-fork servers, freezing in the parent before forking also avoids copy-on-write of pages the collector would otherwise touch in every child.
Verify: after freezing, gc.get_freeze_count() is in the order of your start-up heap, and full-collection pauses disappear from the metric.
3. Do not just raise thresholds¶
The thresholds decide how often each generation is collected. Raising them makes collections rarer — and each young collection bigger, because more objects accumulate between them:
gc.set_threshold(50_000, 50, 100) # fewer, larger collections
Measured: young collections ran about 2,400 times instead of about 67,000, but their median pause rose from 0.11 ms to 3.2 ms and the worst to 44 ms; loop lag p99 rose from 0.2 ms to 10.7 ms. No full collection happened during the run, but nothing about the setting prevents one — it only postpones it, and when it comes it scans the same million objects. Threshold tuning is a throughput lever for batch jobs, where total GC time matters and latency does not; for a latency-sensitive event loop, it trades many invisible pauses for fewer visible ones.
Verify: after any threshold change, compare loop-lag p99 and maximum, not just total GC time.
4. Create fewer cycles in request code¶
The young-generation work — tens of thousands of collections over 15 s here — exists because request handling created reference cycles that reference counting cannot free. Common sources in async code are exceptions kept in variables (an exception's traceback references the frame that holds it), objects with back-references to parents, and closures stored on the objects they close over:
try:
await call_upstream()
except UpstreamError as exc:
error = exc # frame -> error -> traceback -> frame: a cycle
...
error = None # or: del error, or keep only str(exc)
Each cycle broken by design is garbage that reference counting frees immediately, with no collector involved. Use weakref for back-references, avoid storing exceptions beyond the except block — Python deletes the as name at the end of the block for exactly this reason — and check with gc.set_debug(gc.DEBUG_SAVEALL) in a test, which keeps everything the collector would have freed in gc.garbage so you can see what kinds of objects form cycles. The same habits reduce memory growth, as discussed in Memory & Resource Leaks.
Verify: a request-path test run with DEBUG_SAVEALL produces few or no cyclic objects per request.
5. Measure on the exact interpreter you deploy¶
Collector behaviour differs between versions and builds. In these tests, uv's CPython 3.14.6 and 3.13.14 reported thresholds of (2000, 10, 10) and behaved alike: small young collections and a rare full collection of about 1.2 s. The distribution's CPython 3.14.4 reported (2000, 10, 0) and collected the old generation in increments — yet with this workload it still paused for up to 844–1,003 ms and spent most of the run in the collector. The same code, two builds, two different latency profiles:
import gc, sys
print(sys.version, gc.get_threshold(), gc.get_count())
Log this at start-up and include it with every latency benchmark. Re-run the measurement from step 1 after an interpreter upgrade, as part of the checks in measuring asyncio speedups between Python versions, because a collector change can move tail latency by orders of magnitude without changing average throughput at all.
Verify: start-up logs record the interpreter version and GC thresholds, and GC pause metrics are compared before and after upgrades.
Verification¶
Garbage collection is not hurting loop latency when:
- GC pauses are measured per generation, with maxima exported next to loop lag.
- The start-up heap is frozen after a collection, and full-collection pauses are gone.
- Thresholds are left alone in latency-sensitive services, or changed only with lag measurements.
- Request code creates few cycles, and GC behaviour is re-measured on every interpreter upgrade.
Diagnostic Hook: when the loop-lag probe reports a rare stall of hundreds of milliseconds with no slow callback to blame, check the GC pause metric for the same moment. A full collection scanning a large heap is the most common cause of such stalls in otherwise well-behaved asyncio services — and the easiest one to remove.
Pitfalls & edge cases¶
- Large long-lived heaps left unfrozen. Measured: a 1.21–1.28 s full-collection pause.
- Raising thresholds for latency. Measured: worst young collection 44 ms instead of 3.7 ms.
- Freezing objects that later form cycles. Frozen cycles are never collected.
- Assuming one Python build behaves like another. Two 3.14 builds behaved very differently.
Frequently Asked Questions¶
Can garbage collection block the asyncio event loop?
Yes: it runs on the thread that triggers it. With a million long-lived objects, one full collection paused the loop for 1.21 s on CPython 3.14.6 and 1.28 s on 3.13.14 in testing.
What does gc.freeze() do for an asyncio service?
It moves every currently tracked object into a permanent generation that collections skip. Called after loading caches at start-up, it cut the worst loop stall from 1,210 ms to 1.1 ms.
Should I raise gc thresholds to reduce GC pauses?
Not for latency: thresholds of (50,000, 50, 100) made young collections larger (worst 44 ms) and raised loop-lag p99 from 0.2 ms to 10.7 ms. Use gc.freeze() and fewer cycles instead.
How do I find out whether GC causes my latency spikes?
Time collections with gc.callbacks and compare the pause maxima with loop-lag spikes; matching timestamps identify the collector.
Related¶
- Event Loop Configuration — up to the topic overview.
- Raising file descriptor limits for many connections — another default that only bites at scale.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.