Skip to content

Continuous Profiling of Async Services

Traces show which operations are slow; they do not show which lines of Python the event loop was executing while everything waited. A sampling profiler does: it reads the process's stacks many times a second and counts where time goes, without changing the code. Run continuously, it turns "the service got slower after Tuesday's deploy" into a diff of two flame graphs. Measured with py-spy 0.4.2 on Python 3.14, recording an asyncio service at 100 samples per second had no measurable cost — 76,300 requests in 5 s while profiled against 75,300 without — and its profile attributed 79% of samples to one function, a JSON serializer called in every handler, with 62% of all samples inside the standard library's iterencode underneath it. This guide profiles async services safely, reads their profiles correctly, and runs profiling continuously.

Prerequisites

1. Profile a running process without changing it

py-spy reads the target's memory from outside, so the service needs no code changes and no restart:

py-spy top --pid 12345                                   # live view of the hottest functions
py-spy record --pid 12345 --duration 60 -o profile.svg   # flame graph of one minute
py-spy record --rate 100 --format raw -o profile.txt -- python app.py   # profile from start
py-spy dump --pid 12345                                  # one snapshot of every thread's stack

Measured: 100 Hz sampling for the whole run left throughput unchanged within noise (76,300 against 75,300 requests in 5 s). Attaching to another process needs ptrace permission — on many Linux systems kernel.yama.ptrace_scope=1 (as on the test machine) allows it only for children or with CAP_SYS_PTRACE; in containers, add that capability to a debug sidecar or run the service under py-spy record -- python app.py. Starting the profiler as the parent, as in the third command, sidesteps the permission entirely.

Verify: py-spy dump --pid <pid> prints the event loop thread's current stack.

Requests in 5 s with and without py-spy sampling at 100 Hz 2 horizontal bars comparing not profiled with the others. Requests in 5 s with and without py-spy sampling at 100 Hz not profiled 75,300 py-spy record, 100 Hz 76,300 py-spy 0.4.2, Python 3.14; JSON-heavy handler; 526 samples, 0 errors. Out-of-process sampling costs the target almost nothing.

2. Read async profiles with the loop in mind

In an asyncio service, every stack sits under the event loop: run → run_forever → _run_once → Handle._run → the coroutine being stepped. Where samples land below that tells you what the loop is doing:

_run_once (asyncio/base_events.py)            97% inclusive  <- the loop thread, almost always busy
  Handle._run                                 
    handler (app.py)                          88%            <- coroutine steps
      serialize (app.py)                      79%            <- CPU inside one step
        iterencode (json/encoder.py)          62% self
      checksum (app.py)                        7%
  select (selectors.py)                        0%            <- idle time waiting for I/O

Those are the measured shares. Time under select is the loop waiting for I/O — idle, which is good; time under _run_once but outside select is the loop running code, which is the busy fraction. Here the loop was almost never idle, and four-fifths of its time went into serializing responses: the obvious target, either to cache, to switch to a faster encoder, or to move to a process pool. A sampling profiler sees only CPU on the thread: time a coroutine spends awaiting does not appear under that coroutine, which is the opposite of a trace's view.

Verify: the profile's share under select falls as load rises, matching the loop-busy metric.

3. Compare profiles across deploys

A single profile shows the hottest code; two profiles show what changed. Keep profiles per version and diff them:

# collect a minute from the old and the new version under similar load
py-spy record --pid $OLD --duration 60 --format raw -o old.txt
py-spy record --pid $NEW --duration 60 --format raw -o new.txt

# compare self-time shares per function
python - <<'EOF'
import collections
def shares(path):
    c, total = collections.Counter(), 0
    for line in open(path):
        stack, n = line.rsplit(" ", 1)
        c[stack.split(";")[-1].split(" (")[0]] += int(n); total += int(n)
    return {k: v / total for k, v in c.items()}
old, new = shares("old.txt"), shares("new.txt")
for fn in sorted(set(old) | set(new), key=lambda f: new.get(f, 0) - old.get(f, 0), reverse=True)[:10]:
    print(f"{fn:40} {old.get(fn,0):6.1%} -> {new.get(fn,0):6.1%}")
EOF

Raw (collapsed-stack) output is plain text with one line per unique stack and a count, easy to aggregate as above, to diff, or to feed to flame-graph tools. Comparing shares rather than absolute counts makes profiles from different durations and loads comparable. The function whose share rose most is usually the regression.

Verify: a deliberate slowdown introduced in a test build appears at the top of the diff.

Where the loop's time went, by profile share A grid of 5 rows by 3 columns. Where the loop's time went, by profile share frame share meaning _run_once (inclusive) 97% loop thread running handler (inclusive) 88% coroutine steps serialize (inclusive) 79% the hot spot json iterencode (self) 62% inside serialize select (inclusive) 0% idle: loop never waited 526 samples at 100 Hz over about 5 s, py-spy 0.4.2.

4. Run profiling continuously

Profiling on demand catches problems you already know about. A continuous profiler — Grafana Pyroscope, Parca, Datadog, Google Cloud Profiler and others — samples every instance at a low rate all the time and keeps the history:

# Pyroscope's Python agent (in-process), or run py-spy/eBPF agents out of process
import pyroscope

pyroscope.configure(
    application_name="orders-api",
    server_address="http://pyroscope:4040",
    sample_rate=100,                              # Hz
    tags={"version": os.environ.get("APP_VERSION", "dev"), "region": REGION},
)

Tags such as version and region let you compare deploys and instances in the UI, and many backends can link profiles to traces by span or trace id, so a slow trace opens the profile of that moment. eBPF-based agents profile every process on a node from outside, with no per-service setup. Keep the rate modest (around 100 Hz or less); continuous profiling is about statistics over minutes, not individual stacks.

Verify: the profiling backend shows a flame graph per service and version, updated within minutes of a deploy.

5. Combine profiles with traces and loop metrics

Each tool answers a different question, and slow services usually need all three:

# Triage order for "the service is slow":
# 1. Loop metrics: is the loop busy (CPU) or waiting (I/O)?            -> alerting on loop saturation
# 2. Traces:       which operations are slow, and where do they wait?  -> spans per dependency
# 3. Profiles:     if the loop is busy, which code is it running?      -> flame graph share

If loop busy is low and latency is high, the time is spent waiting — traces show on which dependency. If the loop is busy, profiles show which code to optimize or offload. If one instance is slow, compare its profile with a healthy instance's. Profiles are also the quickest way to find accidental blocking calls: a blocking requests.get or time.sleep appears as samples on the loop thread inside socket or sleep functions, under a coroutine that should have been awaiting — the other half of finding blocking calls with asyncio debug mode.

Verify: for the last incident, the post-mortem cites the loop metrics, the trace and the profile that explained it.

Which tool explains this slowness? A decision on What do loop metrics say with 4 outcomes. Which tool explains this slowness? What do loop metrics say? loop mostly idle traces waiting on dependencies loop busy CPU profile which code runs one instance slow compare instance profiles hot key, bad host slower since a deploy diff profiles by version the regression Loop metrics pick the tool; the tool finds the cause.

Verification

Profiling is effective when:

  • Profiles can be taken without restarts, with permissions arranged in advance.
  • Readers know the loop's frames: select is idle, everything else under _run_once is busy.
  • Profiles are kept per version and diffed after deploys.
  • A continuous profiler runs at a low rate and links to traces where supported.

Diagnostic Hook: track, per service version, the share of samples under the loop's selector. A falling idle share at constant traffic means each request costs more CPU — a regression worth a profile diff even before any latency alert fires.

Pitfalls & edge cases

  • Expecting awaits in a CPU profile. Waiting time does not appear under the awaiting coroutine.
  • No ptrace permission in production. Arrange CAP_SYS_PTRACE or a sidecar before an incident.
  • Comparing absolute counts. Use shares across different durations and loads.
  • High sample rates. Continuous profiling needs statistics, not detail.

Frequently Asked Questions

How do I profile a running asyncio service in production?

Attach a sampling profiler such as py-spy (py-spy record --pid or py-spy top) or run a continuous profiling agent. In testing, py-spy at 100 Hz had no measurable effect on throughput.

Why doesn't my profile show time spent awaiting?

A sampling profiler records what the thread is executing. While a coroutine awaits, the loop runs other code or waits in select; use traces to see waiting time.

How do I read an asyncio flame graph?

Everything runs under the event loop's _run_once. Time under select is idle waiting; time under coroutine steps is CPU work, and the widest frames below them are the hot spots.

Does py-spy work with Python 3.14?

py-spy 0.4.2 profiled a Python 3.14 asyncio service in testing, recording 526 samples with no errors.