Skip to content

Returning Memory to the OS After Load Spikes

A service that handles a burst of large requests often keeps its peak memory long after the burst is over, which looks like a leak on a dashboard and counts against container limits as if it were one. Usually it is not a leak: the objects were freed, but the C allocator did not give the pages back to the operating system. Measured on Python 3.14 on Linux with glibc: building 160,000 strings of 600–4,000 bytes raised resident memory from 21 MiB to 396 MiB; keeping 2% of them and freeing the rest left it at 395 MiB. Calling glibc's malloc_trim(0) brought it down to 57 MiB in about 30 ms. That call, made on the event loop, stalled it for 28 ms; made through asyncio.to_thread, the worst stall was 1.3 ms. Doing the spike's allocations in 16 threads, or limiting malloc arenas with MALLOC_ARENA_MAX=2, made no measurable difference here. Forcing allocations through mmap with MALLOC_MMAP_THRESHOLD_=512 lowered the after-free figure to 265 MiB but raised the peak to 508 MiB and slowed the allocation phase by about 55%. This guide tells retained memory from leaked memory and returns it safely.

Prerequisites

1. Reproduce retained memory

Allocate a burst of medium-sized objects — larger than 512 bytes, so CPython hands them to the C allocator rather than its own small-object allocator — keep a scattered few, and free the rest:

def rss_mib() -> int:
    return int(open("/proc/self/status").read().split("VmRSS:")[1].split()[0]) // 1024

r = random.Random(0)
data = [("x" * r.randint(600, 4000), r.random()) for _ in range(160_000)]
survivors = data[::50]                    # 2% outlive the spike, spread through the heap
del data
gc.collect()
print(rss_mib())

Measured: 21 MiB before, 396 MiB at the peak, 395 MiB after freeing 98% of the objects. The memory was free inside the process — new allocations would reuse it — but the allocator kept the pages, because a few live objects were scattered across them and glibc only returns memory automatically from the top of its heap. To the container's memory accounting, this is indistinguishable from a leak.

Verify: after a load spike, compare tracemalloc's traced total — what Python objects use — with resident memory; a large gap with a flat traced total means retained, not leaked.

160,000 strings of 600-4,000 bytes, 2% kept A grid of 5 rows by 4 columns. 160,000 strings of 600-4,000 bytes, 2% kept setup peak after free after malloc_trim(0) defaults, spike on the loop 396 MiB 395 MiB 57 MiB (~30 ms) spike in 16 threads 396 MiB 396 MiB 58 MiB 16 threads, MALLOC_ARENA_MAX=2 396 MiB 396 MiB 57 MiB MALLOC_MMAP_THRESHOLD_=512 508 MiB 265 MiB 61 MiB; build 0.31 s vs 0.20 s MALLOC_TRIM_THRESHOLD_=131072 396 MiB 395 MiB 57 MiB Python 3.14, glibc on Linux.

2. Return free memory with malloc_trim

glibc's malloc_trim(0) walks the heap and releases free pages back to the kernel, including those in the middle of the heap. It is callable from Python through ctypes:

import ctypes

_libc = ctypes.CDLL("libc.so.6")

def release_free_memory() -> None:
    _libc.malloc_trim(0)

Measured: resident memory fell from 395 MiB to 57 MiB, and the call took about 30 ms. The 2% of surviving objects stayed where they were; only empty pages went back. Setting MALLOC_TRIM_THRESHOLD_ lower did nothing here — 395 MiB after freeing — because that threshold governs only the top of the heap, and the free space was in the middle.

Verify: after a spike, a call to malloc_trim(0) lowers resident memory close to what tracemalloc reports in use.

3. Call it off the event loop

malloc_trim is a blocking C call proportional to heap size. On the event loop it stops everything else; ctypes releases the GIL during foreign calls, so a thread avoids that:

async def trim_memory():
    await asyncio.to_thread(release_free_memory)

Measured with a probe timing 1 ms sleeps: calling it directly on the loop produced a worst stall of 27.9–28.4 ms; through asyncio.to_thread, 1.2–1.3 ms, with the same drop from 380 MiB to 42 MiB. Run it after known spikes — the end of a batch import, a large report — or periodically when resident memory exceeds what traced memory explains, rather than after every request.

async def trim_when_bloated(interval=60.0, slack_mib=200):
    while True:
        await asyncio.sleep(interval)
        traced, _ = tracemalloc.get_traced_memory() if tracemalloc.is_tracing() else (0, 0)
        if rss_mib() - traced // 2**20 > slack_mib:
            await trim_memory()

Verify: trimming never stalls the event loop for more than a few milliseconds, and resident memory returns toward baseline after spikes.

Worst event loop stall during malloc_trim(0) 2 horizontal bars comparing on the event loop with the others. Worst event loop stall during malloc_trim(0) on the event loop 28.4 ms via asyncio.to_thread 1.3 ms Same result either way: 380 to 42 MiB.

4. Know what the arena and mmap settings do

Two environment variables are often suggested for this problem. MALLOC_ARENA_MAX limits how many separate heaps glibc creates for threads. Measured with the spike performed in 16 threads, with and without MALLOC_ARENA_MAX=2: identical — 396 MiB at the peak and after freeing, 57–58 MiB after trimming. In CPython most allocation happens while holding the GIL, so threads rarely allocate concurrently enough to need separate arenas; arena limits matter more for C extensions that allocate heavily with the GIL released. Measure before relying on it.

MALLOC_ARENA_MAX=2 python service.py             # measure: no difference in this workload
MALLOC_MMAP_THRESHOLD_=512 python service.py     # every allocation over 512 B is its own mmap

MALLOC_MMAP_THRESHOLD_ makes larger allocations use separate mmap regions, returned to the kernel as soon as they are freed. Measured at 512 bytes: after freeing, 265 MiB instead of 395 — but the peak rose from 396 to 508 MiB, since each small mmap rounds up to a page, and building the data took 0.31 s instead of 0.20 s. A targeted malloc_trim after spikes gave a better result without slowing every allocation.

Verify: any allocator environment variable in production is justified by a measurement on the service's own workload.

5. Tell retention from a leak

Retained memory is free inside the process; leaked memory is held by live objects. The distinction decides the fix:

tracemalloc.start()
await handle_spike()
gc.collect()
traced_mib = tracemalloc.get_traced_memory()[0] / 2**20
print(f"traced {traced_mib:.0f} MiB, resident {rss_mib()} MiB")
await trim_memory()
print(f"after trim: resident {rss_mib()} MiB")

If traced memory stays high, objects are still referenced — a cache, a growing list, exceptions holding frames — and trimming will not help; follow leaks from exceptions holding frames and the other leak guides. If traced memory is low and resident memory falls after a trim, it was retention. Track both in metrics, so the next spike is diagnosed from the dashboard.

Verify: dashboards show resident memory and traced Python memory side by side, and a gap that closes after trimming is labelled as retention.

Is it a leak or retained memory? A flow of 5 stages. Is it a leak or retained memory? After the spike gc.collect() Compare tracemalloc traced vs resident Traced high live objects: find the reference Traced low, resident high malloc_trim in a thread Resident falls retention, not a leak Measured: 395 MiB resident fell to 57 MiB after one trim.

Verification

Memory after spikes is under control when:

  • Retention is distinguished from leaks by comparing traced and resident memory.
  • malloc_trim(0) runs off the event loop, after spikes or when the gap exceeds a threshold.
  • Allocator environment variables are used only where measured to help.
  • Dashboards show both measures, so retention does not trigger leak investigations.

Diagnostic Hook: when a service's memory stays at its peak after traffic drops, call malloc_trim(0) once through asyncio.to_thread and watch resident memory. A drop like 395 MiB to 57 MiB means the allocator was holding freed memory, not that objects leaked.

Pitfalls & edge cases

  • Treating retained memory as a leak. Measured: 98% of objects freed, resident memory unchanged.
  • Calling malloc_trim on the loop. Measured: a 28 ms stall.
  • Expecting MALLOC_ARENA_MAX to fix it. No difference here.
  • Lowering the mmap threshold globally. Measured: higher peak, 55% slower allocation.

Frequently Asked Questions

Why doesn't my Python service release memory after a spike?

glibc keeps freed memory inside the process when live objects are scattered among free pages. With 2% of objects kept, resident memory stayed at 395 of a 396 MiB peak.

How do I make Python return memory to the OS?

On glibc, call malloc_trim(0) through ctypes. It took resident memory from 395 MiB to 57 MiB in about 30 ms.

Is calling malloc_trim safe in an asyncio service?

Yes, but run it with asyncio.to_thread: on the loop it caused a 28 ms stall; in a thread the worst stall was 1.3 ms.

Does MALLOC_ARENA_MAX reduce Python memory usage?

Not in this workload: the spike in 16 threads behaved the same with MALLOC_ARENA_MAX=2. It matters most for C code allocating without the GIL.