Continuous Profiling of Async Services¶
Traces show which operations are slow; they do not show which lines of Python the event loop was executing while everything waited. A sampling profiler does: it reads the process's stacks many times a second and counts where time goes, without changing the code. Run continuously, it turns "the service got slower after Tuesday's deploy" into a diff of two flame graphs. Measured with py-spy 0.4.2 on Python 3.14, recording an asyncio service at 100 samples per second had no measurable cost — 76,300 requests in 5 s while profiled against 75,300 without — and its profile attributed 79% of samples to one function, a JSON serializer called in every handler, with 62% of all samples inside the standard library's iterencode underneath it. This guide profiles async services safely, reads their profiles correctly, and runs profiling continuously.
Prerequisites¶
- Python 3.11+,
pip install py-spy(measured with 0.4.2); Linux needs ptrace permission to attach to a running process. - Loop saturation signals, from alerting on event loop saturation.
- Tracing, from tracing asyncio services with OpenTelemetry.
1. Profile a running process without changing it¶
py-spy reads the target's memory from outside, so the service needs no code changes and no restart:
py-spy top --pid 12345 # live view of the hottest functions
py-spy record --pid 12345 --duration 60 -o profile.svg # flame graph of one minute
py-spy record --rate 100 --format raw -o profile.txt -- python app.py # profile from start
py-spy dump --pid 12345 # one snapshot of every thread's stack
Measured: 100 Hz sampling for the whole run left throughput unchanged within noise (76,300 against 75,300 requests in 5 s). Attaching to another process needs ptrace permission — on many Linux systems kernel.yama.ptrace_scope=1 (as on the test machine) allows it only for children or with CAP_SYS_PTRACE; in containers, add that capability to a debug sidecar or run the service under py-spy record -- python app.py. Starting the profiler as the parent, as in the third command, sidesteps the permission entirely.
Verify: py-spy dump --pid <pid> prints the event loop thread's current stack.
2. Read async profiles with the loop in mind¶
In an asyncio service, every stack sits under the event loop: run → run_forever → _run_once → Handle._run → the coroutine being stepped. Where samples land below that tells you what the loop is doing:
_run_once (asyncio/base_events.py) 97% inclusive <- the loop thread, almost always busy
Handle._run
handler (app.py) 88% <- coroutine steps
serialize (app.py) 79% <- CPU inside one step
iterencode (json/encoder.py) 62% self
checksum (app.py) 7%
select (selectors.py) 0% <- idle time waiting for I/O
Those are the measured shares. Time under select is the loop waiting for I/O — idle, which is good; time under _run_once but outside select is the loop running code, which is the busy fraction. Here the loop was almost never idle, and four-fifths of its time went into serializing responses: the obvious target, either to cache, to switch to a faster encoder, or to move to a process pool. A sampling profiler sees only CPU on the thread: time a coroutine spends awaiting does not appear under that coroutine, which is the opposite of a trace's view.
Verify: the profile's share under select falls as load rises, matching the loop-busy metric.
3. Compare profiles across deploys¶
A single profile shows the hottest code; two profiles show what changed. Keep profiles per version and diff them:
# collect a minute from the old and the new version under similar load
py-spy record --pid $OLD --duration 60 --format raw -o old.txt
py-spy record --pid $NEW --duration 60 --format raw -o new.txt
# compare self-time shares per function
python - <<'EOF'
import collections
def shares(path):
c, total = collections.Counter(), 0
for line in open(path):
stack, n = line.rsplit(" ", 1)
c[stack.split(";")[-1].split(" (")[0]] += int(n); total += int(n)
return {k: v / total for k, v in c.items()}
old, new = shares("old.txt"), shares("new.txt")
for fn in sorted(set(old) | set(new), key=lambda f: new.get(f, 0) - old.get(f, 0), reverse=True)[:10]:
print(f"{fn:40} {old.get(fn,0):6.1%} -> {new.get(fn,0):6.1%}")
EOF
Raw (collapsed-stack) output is plain text with one line per unique stack and a count, easy to aggregate as above, to diff, or to feed to flame-graph tools. Comparing shares rather than absolute counts makes profiles from different durations and loads comparable. The function whose share rose most is usually the regression.
Verify: a deliberate slowdown introduced in a test build appears at the top of the diff.
4. Run profiling continuously¶
Profiling on demand catches problems you already know about. A continuous profiler — Grafana Pyroscope, Parca, Datadog, Google Cloud Profiler and others — samples every instance at a low rate all the time and keeps the history:
# Pyroscope's Python agent (in-process), or run py-spy/eBPF agents out of process
import pyroscope
pyroscope.configure(
application_name="orders-api",
server_address="http://pyroscope:4040",
sample_rate=100, # Hz
tags={"version": os.environ.get("APP_VERSION", "dev"), "region": REGION},
)
Tags such as version and region let you compare deploys and instances in the UI, and many backends can link profiles to traces by span or trace id, so a slow trace opens the profile of that moment. eBPF-based agents profile every process on a node from outside, with no per-service setup. Keep the rate modest (around 100 Hz or less); continuous profiling is about statistics over minutes, not individual stacks.
Verify: the profiling backend shows a flame graph per service and version, updated within minutes of a deploy.
5. Combine profiles with traces and loop metrics¶
Each tool answers a different question, and slow services usually need all three:
# Triage order for "the service is slow":
# 1. Loop metrics: is the loop busy (CPU) or waiting (I/O)? -> alerting on loop saturation
# 2. Traces: which operations are slow, and where do they wait? -> spans per dependency
# 3. Profiles: if the loop is busy, which code is it running? -> flame graph share
If loop busy is low and latency is high, the time is spent waiting — traces show on which dependency. If the loop is busy, profiles show which code to optimize or offload. If one instance is slow, compare its profile with a healthy instance's. Profiles are also the quickest way to find accidental blocking calls: a blocking requests.get or time.sleep appears as samples on the loop thread inside socket or sleep functions, under a coroutine that should have been awaiting — the other half of finding blocking calls with asyncio debug mode.
Verify: for the last incident, the post-mortem cites the loop metrics, the trace and the profile that explained it.
Verification¶
Profiling is effective when:
- Profiles can be taken without restarts, with permissions arranged in advance.
- Readers know the loop's frames:
selectis idle, everything else under_run_onceis busy. - Profiles are kept per version and diffed after deploys.
- A continuous profiler runs at a low rate and links to traces where supported.
Diagnostic Hook: track, per service version, the share of samples under the loop's selector. A falling idle share at constant traffic means each request costs more CPU — a regression worth a profile diff even before any latency alert fires.
Pitfalls & edge cases¶
- Expecting awaits in a CPU profile. Waiting time does not appear under the awaiting coroutine.
- No ptrace permission in production. Arrange
CAP_SYS_PTRACEor a sidecar before an incident. - Comparing absolute counts. Use shares across different durations and loads.
- High sample rates. Continuous profiling needs statistics, not detail.
Frequently Asked Questions¶
How do I profile a running asyncio service in production?
Attach a sampling profiler such as py-spy (py-spy record --pid
Why doesn't my profile show time spent awaiting?
A sampling profiler records what the thread is executing. While a coroutine awaits, the loop runs other code or waits in select; use traces to see waiting time.
How do I read an asyncio flame graph?
Everything runs under the event loop's _run_once. Time under select is idle waiting; time under coroutine steps is CPU work, and the widest frames below them are the hot spots.
Does py-spy work with Python 3.14?
py-spy 0.4.2 profiled a Python 3.14 asyncio service in testing, recording 526 samples with no errors.
Related¶
- Observability & Tracing — up to the topic overview.
- Profiling async services with memray — the memory counterpart.
- Resilience, Cancellation & Error Handling — the section overview.