Skip to content

Catching Leaks in CI with Memory Budgets

Most memory leaks in async services are found in production, days after the deploy that introduced them, because nothing in the test suite runs a code path thousands of times and looks at memory afterwards. A memory budget test does exactly that, and it is cheap. Measured on Python 3.14 with pytest 9: a test that warmed up a request handler with 200 calls, then ran it 2,000 more times between two tracemalloc snapshots and asserted growth per call and task count, took 0.36 s. Against the clean handler it measured 0.4 bytes of growth per call. Three seeded leaks — an unbounded cache, a background task created per request, and a callback registered per request — grew 1,322, 1,202 and 1,487 bytes per call and all failed a 64-byte budget, the task leak also failing the task-count check with 2,000 extra tasks. A fourth, smaller leak — one integer appended per request — grew 40.6–41.8 bytes per call and passed the 64-byte budget. Since the clean handler's growth was 0.3–0.4 bytes per call across five runs, a budget of 8 bytes would have caught it. This guide builds the test and sets its budget from measured noise.

Prerequisites

1. Write the budget test

Run the code path enough times that a per-call leak dominates any one-off allocation, with a warm-up so caches and lazy imports are already full:

WARMUP, N = 200, 2000
BUDGET_BYTES_PER_CALL = 8

def test_handle_memory_budget():
    async def scenario():
        for i in range(WARMUP):
            await handle(i)
        gc.collect()
        tracemalloc.start()
        before = tracemalloc.take_snapshot()
        tasks_before = len(asyncio.all_tasks())

        for i in range(WARMUP, WARMUP + N):
            await handle(i)

        gc.collect()
        after = tracemalloc.take_snapshot()
        tracemalloc.stop()
        diff = after.compare_to(before, "lineno")
        per_call = sum(s.size_diff for s in diff) / N
        assert len(asyncio.all_tasks()) <= tasks_before, "tasks leaked"
        assert per_call < BUDGET_BYTES_PER_CALL, f"{per_call:.1f} B/call; top: {diff[0]}"
    asyncio.run(scenario())

Measured against the clean handler: 0.4 bytes of growth per call, and the test took 0.36 s including pytest start-up. The task-count check matters as much as the memory check: a leaked background task holds memory, but it also keeps running.

Verify: the test passes against current code and finishes in well under a second.

Budget test against seeded leaks, 2,000 calls after 200 warm-up A grid of 5 rows by 5 columns. Budget test against seeded leaks, 2,000 calls after 200 warm-up handler growth per call extra tasks 64 B budget 8 B budget clean 0.4 B 0 pass pass unbounded cache keyed by request 1,322 B 0 fail fail background task per request 1,202 B 2,000 fail fail callback registered per request 1,487 B 0 fail fail one int appended per request 40.6 B 0 pass (missed) fail Python 3.14, pytest 9; each run took 0.33-0.40 s.

2. Point the failure at the leak

A budget failure should say where the memory went. compare_to(..., "lineno") sorts by growth, so the first entry is usually the culprit:

diff = after.compare_to(before, "lineno")
for stat in diff[:3]:
    print(stat)       # file:line, total size, change, count

Measured: for the task leak, the top entry was the create_task line in the handler, with 2,001 new allocations; for the cache and callback leaks, it was json/encoder.py, where the leaked strings were created — the allocation site, not the place that kept them. Allocation sites answer "what is leaking", and the code that stores the object is usually one call up; tracemalloc.start(10) records deeper tracebacks for each allocation, at a higher cost, when the top frame is not enough.

Verify: a failing budget test prints the top growth sites, and the leak can be located from that output alone.

3. Set the budget from measured noise

The budget is a threshold between noise and leak. Too high, and small leaks pass; too low, and the test is flaky. Measure the clean path's growth across several runs:

async def growth_per_call(n=2000):
    for i in range(200):
        await handle(i)
    gc.collect(); tracemalloc.start(); s0 = tracemalloc.take_snapshot()
    for i in range(200, 200 + n):
        await handle(i)
    gc.collect(); s1 = tracemalloc.take_snapshot()
    for i in range(200 + n, 200 + 2 * n):
        await handle(i)
    gc.collect(); s2 = tracemalloc.take_snapshot(); tracemalloc.stop()
    first = sum(s.size_diff for s in s1.compare_to(s0, "filename")) / n
    second = sum(s.size_diff for s in s2.compare_to(s1, "filename")) / n
    return first, second

Measured over five runs of the clean handler: 0.4 bytes per call in the first window and 0.3 in the second, every time. The small leak measured 40.6 and 41.8 — the same in both windows, as a true per-call leak is. A budget of 8 bytes, twenty times the noise, separates them; the earlier 64-byte budget let the small leak through. One integer per request sounds harmless, but at 1,000 requests per second it is 3.5 GB a day.

Verify: the budget sits well above the clean path's measured growth and well below the smallest leak the team cares about.

Growth per call, bytes 3 horizontal bars comparing clean handler (noise) with the others. Growth per call, bytes clean handler (noise) 0.4 B one int per request 40.6 B unbounded cache 1,322 B An 8 B budget catches all three leaks; a 64 B budget misses the middle one. Same growth in two consecutive windows means a real per-call leak.

4. Budget other resources the same way

Memory is one leaked resource among several. The same before-and-after pattern catches the others:

def open_fds() -> int:
    return len(os.listdir("/proc/self/fd"))

fds_before, tasks_before = open_fds(), len(asyncio.all_tasks())
for i in range(N):
    await handle(i)
await asyncio.sleep(0)                                   # let finished tasks be collected
assert open_fds() <= fds_before, "file descriptors leaked"
assert len(asyncio.all_tasks()) <= tasks_before, "tasks leaked"

Measured, the task check alone flagged the background-task leak with 2,000 extra tasks. File descriptors catch unclosed sockets and files, the usual partner of an unclosed client session; see detecting leaked sockets and file descriptors. Use real dependencies or faithful fakes in these tests — a mocked HTTP client cannot leak connections.

Verify: budget tests assert on memory, tasks and file descriptors for each long-running code path.

5. Run budgets in CI on the paths that matter

Budget tests are fast, so they can run on every change. Choose the code paths that run per request or per message — handlers, consumers, middleware — rather than start-up code, and mark them so they run together:

@pytest.mark.memory_budget
def test_order_handler_budget(): ...

@pytest.mark.memory_budget
def test_kafka_consumer_budget(): ...
pytest -m memory_budget -p no:randomly      # stable order: other tests' allocations stay out

Keep each budget test in its own process if test order affects the numbers — pytest-forked or a separate CI step — since other tests can leave caches warming up during the measured window. When a budget test fails after a dependency upgrade, the top allocation sites show whether the leak is in your code or the library's. For leaks that only appear under concurrency, the same measurement wrapped around a concurrent workload is in tracking task growth in long-running services.

Verify: the budget suite runs on every change and has caught at least one seeded leak per resource type.

A memory budget test A flow of 5 stages. A memory budget test Warm up 200 calls, caches full Snapshot tracemalloc, tasks, fds Exercise 2,000 calls Compare bytes per call vs budget On failure top allocation sites 0.36 s per path; budget set at 20x the clean path's noise.

Verification

Leak budgets work when:

  • Each per-request code path has a budget test with warm-up, snapshots and a per-call threshold.
  • The budget is set from measured noise, low enough to catch small leaks.
  • Tasks and file descriptors are budgeted alongside memory.
  • Failures print allocation sites, and seeded leaks of each kind are caught.

Diagnostic Hook: when a memory budget test fails intermittently, measure the clean path's growth per call in two consecutive windows. Noise varies between windows and runs; a real leak grows by the same amount per call in both, as the 40.6 and 41.8 bytes per call of the small leak did here.

Pitfalls & edge cases

  • Budgets set by guesswork. Measured: a 64 B budget missed a 41 B/call leak.
  • No warm-up. First-call caches look like leaks.
  • Mocked dependencies. They cannot leak what real ones do.
  • Memory checks without task checks. Leaked tasks keep running, not just using memory.

Frequently Asked Questions

How do I test for memory leaks in pytest?

Warm up the code path, take a tracemalloc snapshot, run it 2,000 times, snapshot again and assert growth per call is under a budget. It took 0.36 s.

What budget should a memory leak test use?

Measure the clean path first: it grew 0.3-0.4 bytes per call here, so 8 bytes caught a 41-byte leak that a 64-byte budget missed.

How do I find what is leaking from a failing budget test?

Print the top entries of after.compare_to(before, "lineno"). It pointed at the create_task line for a task leak; for cached data it showed where the data was created.

Should leak tests check asyncio tasks too?

Yes. A task created per request added 2,000 tasks over 2,000 calls, caught by comparing len(asyncio.all_tasks()) before and after.