Returning Memory to the OS After Load Spikes¶
A service that handles a burst of large requests often keeps its peak memory long after the burst is over, which looks like a leak on a dashboard and counts against container limits as if it were one. Usually it is not a leak: the objects were freed, but the C allocator did not give the pages back to the operating system. Measured on Python 3.14 on Linux with glibc: building 160,000 strings of 600–4,000 bytes raised resident memory from 21 MiB to 396 MiB; keeping 2% of them and freeing the rest left it at 395 MiB. Calling glibc's malloc_trim(0) brought it down to 57 MiB in about 30 ms. That call, made on the event loop, stalled it for 28 ms; made through asyncio.to_thread, the worst stall was 1.3 ms. Doing the spike's allocations in 16 threads, or limiting malloc arenas with MALLOC_ARENA_MAX=2, made no measurable difference here. Forcing allocations through mmap with MALLOC_MMAP_THRESHOLD_=512 lowered the after-free figure to 265 MiB but raised the peak to 508 MiB and slowed the allocation phase by about 55%. This guide tells retained memory from leaked memory and returns it safely.
Prerequisites¶
- Python 3.11+ on Linux with glibc;
malloc_trimis glibc-specific (musl-based images such as Alpine behave differently). - Leak hunting, from finding memory leaks in asyncio with tracemalloc.
- The topic overview, Memory & Resource Leaks.
1. Reproduce retained memory¶
Allocate a burst of medium-sized objects — larger than 512 bytes, so CPython hands them to the C allocator rather than its own small-object allocator — keep a scattered few, and free the rest:
def rss_mib() -> int:
return int(open("/proc/self/status").read().split("VmRSS:")[1].split()[0]) // 1024
r = random.Random(0)
data = [("x" * r.randint(600, 4000), r.random()) for _ in range(160_000)]
survivors = data[::50] # 2% outlive the spike, spread through the heap
del data
gc.collect()
print(rss_mib())
Measured: 21 MiB before, 396 MiB at the peak, 395 MiB after freeing 98% of the objects. The memory was free inside the process — new allocations would reuse it — but the allocator kept the pages, because a few live objects were scattered across them and glibc only returns memory automatically from the top of its heap. To the container's memory accounting, this is indistinguishable from a leak.
Verify: after a load spike, compare tracemalloc's traced total — what Python objects use — with resident memory; a large gap with a flat traced total means retained, not leaked.
2. Return free memory with malloc_trim¶
glibc's malloc_trim(0) walks the heap and releases free pages back to the kernel, including those in the middle of the heap. It is callable from Python through ctypes:
import ctypes
_libc = ctypes.CDLL("libc.so.6")
def release_free_memory() -> None:
_libc.malloc_trim(0)
Measured: resident memory fell from 395 MiB to 57 MiB, and the call took about 30 ms. The 2% of surviving objects stayed where they were; only empty pages went back. Setting MALLOC_TRIM_THRESHOLD_ lower did nothing here — 395 MiB after freeing — because that threshold governs only the top of the heap, and the free space was in the middle.
Verify: after a spike, a call to malloc_trim(0) lowers resident memory close to what tracemalloc reports in use.
3. Call it off the event loop¶
malloc_trim is a blocking C call proportional to heap size. On the event loop it stops everything else; ctypes releases the GIL during foreign calls, so a thread avoids that:
async def trim_memory():
await asyncio.to_thread(release_free_memory)
Measured with a probe timing 1 ms sleeps: calling it directly on the loop produced a worst stall of 27.9–28.4 ms; through asyncio.to_thread, 1.2–1.3 ms, with the same drop from 380 MiB to 42 MiB. Run it after known spikes — the end of a batch import, a large report — or periodically when resident memory exceeds what traced memory explains, rather than after every request.
async def trim_when_bloated(interval=60.0, slack_mib=200):
while True:
await asyncio.sleep(interval)
traced, _ = tracemalloc.get_traced_memory() if tracemalloc.is_tracing() else (0, 0)
if rss_mib() - traced // 2**20 > slack_mib:
await trim_memory()
Verify: trimming never stalls the event loop for more than a few milliseconds, and resident memory returns toward baseline after spikes.
4. Know what the arena and mmap settings do¶
Two environment variables are often suggested for this problem. MALLOC_ARENA_MAX limits how many separate heaps glibc creates for threads. Measured with the spike performed in 16 threads, with and without MALLOC_ARENA_MAX=2: identical — 396 MiB at the peak and after freeing, 57–58 MiB after trimming. In CPython most allocation happens while holding the GIL, so threads rarely allocate concurrently enough to need separate arenas; arena limits matter more for C extensions that allocate heavily with the GIL released. Measure before relying on it.
MALLOC_ARENA_MAX=2 python service.py # measure: no difference in this workload
MALLOC_MMAP_THRESHOLD_=512 python service.py # every allocation over 512 B is its own mmap
MALLOC_MMAP_THRESHOLD_ makes larger allocations use separate mmap regions, returned to the kernel as soon as they are freed. Measured at 512 bytes: after freeing, 265 MiB instead of 395 — but the peak rose from 396 to 508 MiB, since each small mmap rounds up to a page, and building the data took 0.31 s instead of 0.20 s. A targeted malloc_trim after spikes gave a better result without slowing every allocation.
Verify: any allocator environment variable in production is justified by a measurement on the service's own workload.
5. Tell retention from a leak¶
Retained memory is free inside the process; leaked memory is held by live objects. The distinction decides the fix:
tracemalloc.start()
await handle_spike()
gc.collect()
traced_mib = tracemalloc.get_traced_memory()[0] / 2**20
print(f"traced {traced_mib:.0f} MiB, resident {rss_mib()} MiB")
await trim_memory()
print(f"after trim: resident {rss_mib()} MiB")
If traced memory stays high, objects are still referenced — a cache, a growing list, exceptions holding frames — and trimming will not help; follow leaks from exceptions holding frames and the other leak guides. If traced memory is low and resident memory falls after a trim, it was retention. Track both in metrics, so the next spike is diagnosed from the dashboard.
Verify: dashboards show resident memory and traced Python memory side by side, and a gap that closes after trimming is labelled as retention.
Verification¶
Memory after spikes is under control when:
- Retention is distinguished from leaks by comparing traced and resident memory.
malloc_trim(0)runs off the event loop, after spikes or when the gap exceeds a threshold.- Allocator environment variables are used only where measured to help.
- Dashboards show both measures, so retention does not trigger leak investigations.
Diagnostic Hook: when a service's memory stays at its peak after traffic drops, call malloc_trim(0) once through asyncio.to_thread and watch resident memory. A drop like 395 MiB to 57 MiB means the allocator was holding freed memory, not that objects leaked.
Pitfalls & edge cases¶
- Treating retained memory as a leak. Measured: 98% of objects freed, resident memory unchanged.
- Calling
malloc_trimon the loop. Measured: a 28 ms stall. - Expecting
MALLOC_ARENA_MAXto fix it. No difference here. - Lowering the mmap threshold globally. Measured: higher peak, 55% slower allocation.
Frequently Asked Questions¶
Why doesn't my Python service release memory after a spike?
glibc keeps freed memory inside the process when live objects are scattered among free pages. With 2% of objects kept, resident memory stayed at 395 of a 396 MiB peak.
How do I make Python return memory to the OS?
On glibc, call malloc_trim(0) through ctypes. It took resident memory from 395 MiB to 57 MiB in about 30 ms.
Is calling malloc_trim safe in an asyncio service?
Yes, but run it with asyncio.to_thread: on the loop it caused a 28 ms stall; in a thread the worst stall was 1.3 ms.
Does MALLOC_ARENA_MAX reduce Python memory usage?
Not in this workload: the spike in 16 threads behaved the same with MALLOC_ARENA_MAX=2. It matters most for C code allocating without the GIL.
Related¶
- Memory & Resource Leaks — up to the topic overview.
- Measuring memory per connection — the steady-state side of memory planning.
- Resilience, Cancellation & Error Handling — the section overview.