Skip to content

Profiling Async Services with memray

tracemalloc tells you which Python lines allocate; memray records every allocation — Python and native — with its full stack, and then answers questions about the recording afterwards: what was alive at the peak, what was never freed, which call paths allocate the most. Profiled on Python 3.14 with memray 1.20, a small asyncio service that rendered JSON per request and accidentally kept every body in a module-level dict handled 131,500 requests in 3 s under memray against 179,000 without it — about 27% slower — and wrote a 0.5 MiB capture file for 941,723 allocations. Its leak report attributed 197.8 MiB of the 202.8 MiB never freed to a single line, the json.dumps whose result went into the cache; the allocation statistics separately showed 1.05 GB of short-lived allocations in the JSON encoder that were not the problem. This guide captures a service, reads the reports that matter for async code, and keeps the overhead acceptable.

Prerequisites

1. Capture the service under realistic load

Run the service under memray run, drive it with load for long enough to show the problem, and stop it cleanly:

python -m memray run -o service.bin -m myservice           # or: memray run -o service.bin app.py
# drive traffic, then stop the service with SIGINT/SIGTERM so the capture is finalized

python -m memray run --native -o service.bin -m myservice  # include C-extension frames
python -m memray run --follow-fork -o service.bin -m myservice   # pre-fork servers

Measured overhead on an allocation-heavy service: 27% fewer requests in the same time, a 0.5 MiB capture for 3 seconds. Allocation-light services see less; native tracking (--native) costs more but shows allocations inside C extensions such as parsers and drivers. For servers that fork workers, --follow-fork gives one capture per process. Run captures in staging or on one canary instance, not across the fleet.

Verify: the capture file exists and memray stats service.bin reports allocation totals that match the traffic you sent.

The cost of capturing with memray 2 horizontal bars comparing requests in 3 s, plain with the others. The cost of capturing with memray requests in 3 s, plain 179,000 requests in 3 s, under memray 131,500 (-27%) memray 1.20, Python 3.14; a JSON-rendering service with about 300,000 allocations per second; capture file 0.5 MiB. Affordable on a canary or in staging, not on every instance.

2. Read the leak report first

For a service whose memory grows, the question is "what was allocated and never freed". memray's leak view answers it directly:

python -m memray flamegraph --leaks -o leaks.html service.bin    # interactive, by call stack
python -m memray summary service.bin                             # terminal table, by location
import collections
import memray

reader = memray.FileReader("service.bin")
leaked = collections.Counter()
for record in reader.get_leaked_allocation_records(merge_threads=True):
    frame = next((f for f in record.stack_trace() if "myservice" in f[1]), record.stack_trace()[0])
    leaked[f"{frame[0]} {frame[1]}:{frame[2]}"] += record.size
for location, size in leaked.most_common(5):
    print(f"{size / 2**20:8.1f} MiB  {location}")

Measured: 202.8 MiB was never freed by the end of the run, 197.8 MiB of it allocated at render svc.py:5 — the json.dumps call whose result was stored in the cache on the next line — and 5.0 MiB by the cache dict itself growing. Attributing leaked bytes to the first frame in your own code, as the script does, skips over the library frames (the JSON encoder) where the bytes were technically allocated. In async code, stacks pass through the event loop's _run_once and task machinery; the frames that matter are the coroutine frames below them.

Verify: the top leak location points at code that stores objects beyond a request's lifetime.

3. Separate churn from retention

Total allocated bytes and allocation counts measure churn — work the allocator does — not memory held. They are a different problem:

python -m memray stats service.bin
#   Total memory allocated: 1.081GB      <- churn over 3 s
#   Peak memory usage:      212.642MB    <- what was held at once
#   Top location by size: iterencode json/encoder.py:263 -> 1.046GB

Measured: the JSON encoder allocated 1.05 GB in 920,514 allocations over the run — almost all freed immediately — while the leak was 0.2 GB. Churn matters for CPU (allocation and freeing are not free) and for fragmentation, and the stats view is how to find it; retention is what makes memory grow, and the leak view is how to find that. Do not chase the biggest allocator in stats when the symptom is growth.

Verify: the peak memory in memray stats roughly matches the process's resident memory at the end of the load.

Which memray view answers which question A grid of 4 rows by 3 columns. Which memray view answers which question question view in the test what grows and is never freed? flamegraph --leaks / leaked records 197.8 MiB at render svc.py:5 what allocates the most (churn)? stats, summary encoder: 1.05 GB, freed what was alive at the peak? flamegraph (default) 212.6 MiB peak when did memory grow? flamegraph --temporal steady climb Retention and churn are different problems with different views.

4. Watch memory over time with live and temporal views

Some problems are not leaks but spikes: one request type that briefly needs a gigabyte, a batch that loads everything at once. The temporal flamegraph shows memory over the run and what was alive at any selected moment:

python -m memray flamegraph --temporal -o temporal.html service.bin

python -m memray run --live-remote --live-port 12345 -m myservice     # in one terminal
python -m memray live 12345                                           # in another: live TUI

The temporal view distinguishes a steady climb (retention) from a sawtooth (spikes that are freed) and lets you select the peak to see which stacks held memory at that moment. Live mode shows current allocations while the service runs, which is useful for reproducing a spike interactively. memray can also attach to an already-running process (memray attach <pid>), using a debugger under the hood; it needs gdb or lldb and permission to trace the process.

Verify: the temporal view of a captured spike shows which coroutine's frames held the peak allocation.

5. Make memray part of investigating, not monitoring

memray is an investigation tool. Pair it with cheap continuous signals that tell you when to reach for it:

# continuous, cheap: export these and alert on trends
import gc
import resource

def memory_metrics() -> dict:
    return {
        "rss_mib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024,   # Linux: KiB
        "gc_objects": len(gc.get_objects()),
        "tasks": len(asyncio.all_tasks()),
    }

Resident memory, live object count and task count, exported every few seconds, show that something grows and roughly what kind of thing; a memray capture on a canary under replayed traffic then shows where. Keep a capture from a healthy version to compare with — the difference between two leak reports is often the fastest way to the regression. Delete captures afterwards; they contain stack data from your service.

Verify: a dashboard shows RSS, object count and task count per instance, and the team knows how to run a memray capture on a canary.

Which memory tool for this symptom? A decision on What does memory do with 4 outcomes. Which memory tool for this symptom? What does memory do? grows steadily memray run + leak report top retaining line allocation burns CPU memray stats churn by location spikes and recovers temporal flamegraph / live what held the peak always, in production RSS, gc objects, tasks cheap trend metrics Monitor cheaply everywhere; capture deeply in one place.

Verification

memray is used effectively when:

  • Captures run on a canary or in staging under realistic load, and are finalized cleanly.
  • Leak reports are read first for growth, attributed to your own frames.
  • Churn and retention are treated separately.
  • Cheap metrics trigger captures, and a healthy baseline capture exists for comparison.

Diagnostic Hook: after a fix, rerun the same load under memray and compare total leaked bytes with the previous capture. The leaked total should fall to near zero for a steady-state service; if it only shrinks, the report's next location is the next retention site.

Pitfalls & edge cases

  • Capturing in production at scale. Measured: 27% throughput cost.
  • Chasing churn when memory grows. The encoder's 1.05 GB was freed; the leak was elsewhere.
  • Killing the process before the capture is finalized. Stop it with a signal it handles.
  • Reading library frames as the cause. Attribute to the first frame in your code.

Frequently Asked Questions

How do I find a memory leak in an asyncio service with memray?

Run the service with python -m memray run -o out.bin, drive realistic load, stop it cleanly, and open memray flamegraph --leaks or iterate FileReader.get_leaked_allocation_records(). In testing the report attributed 197.8 of 202.8 MiB to one line.

How much overhead does memray add?

It depends on allocation rate; an allocation-heavy asyncio service handled 27% fewer requests under memray in testing. Native tracking adds more.

What is the difference between memray stats and the leaks report?

stats shows how much was allocated in total, mostly short-lived churn; the leaks report shows what was still allocated at the end, which is what makes memory grow.

Can memray profile a running Python process?

Yes, with memray attach , which needs gdb or lldb and permission to trace the process; or start the service with --live-remote and connect with memray live.