Microbenchmarking Coroutines with pyperf and timeit¶
Small async operations — creating a task, yielding to the loop, awaiting a coroutine — cost microseconds or less, and their costs decide whether a design that does them thousands of times per request is viable. Measuring them is easy to get wrong because the event loop around them is far more expensive than they are. Measured on Python 3.14 with pyperf 2.10, running many iterations inside one loop across several processes: awaiting a coroutine directly cost 48.6 ns ± 1.2, asyncio.sleep(0) 1.48 µs ± 0.02, create_task plus await 3.15 µs ± 0.06, and gather over 100 coroutines 150 µs (1.5 µs each). A timeit loop that called asyncio.run() for each measurement reported 44.0 µs for the same create-task operation — fourteen times too high, because it was timing loop creation. A timeit run that reused one loop and amortized over 10,000 iterations agreed with pyperf: 3.19 µs median. This guide benchmarks async code so the number describes the code.
Prerequisites¶
- Python 3.11+,
pip install pyperf(tested with 2.10). - What these operations do, from Task Scheduling & Lifecycle.
- A quiet machine: microbenchmarks are sensitive to other load; see step 5.
1. Time the operation inside one long-lived loop¶
The event loop's creation and teardown cost tens of microseconds. Put the loop outside the timed region and many iterations inside it:
import asyncio
import timeit
async def noop():
return 1
async def create_task_and_await():
return await asyncio.create_task(noop())
# Wrong: times asyncio.run - loop creation and teardown - for every call
per_call = timeit.timeit(lambda: asyncio.run(create_task_and_await()), number=2000) / 2000
# measured: 44.0 us
# Better: one loop, the iterations inside the coroutine
async def many(n: int) -> None:
for _ in range(n):
await create_task_and_await()
loop = asyncio.new_event_loop()
runs = [timeit.timeit(lambda: loop.run_until_complete(many(10_000)), number=1) / 10_000
for _ in range(7)]
# measured: median 3.19 us (min 3.18, max 3.27)
The first version measured 44.0 µs, of which about 41 µs was asyncio.run. The second amortizes the single run_until_complete over 10,000 iterations and reports the operation itself. Repeating the measurement and taking the median and spread shows whether the number is stable.
Verify: doubling the iteration count inside the loop leaves the per-iteration figure unchanged; if it falls, setup is still inside the timed region.
2. Use pyperf for numbers you will publish¶
pyperf runs benchmarks in several fresh worker processes, calibrates the iteration count, discards warm-up runs and reports the mean with its standard deviation. For coroutines, bench_async_func runs the coroutine function many times inside one loop per process:
import asyncio
import pyperf
async def noop():
return 1
async def create_task_and_await():
return await asyncio.create_task(noop())
async def sleep0():
await asyncio.sleep(0)
async def gather100():
await asyncio.gather(*(noop() for _ in range(100)))
runner = pyperf.Runner()
runner.bench_async_func("await coroutine", noop)
runner.bench_async_func("create_task + await", create_task_and_await)
runner.bench_async_func("asyncio.sleep(0)", sleep0)
runner.bench_async_func("gather of 100 coroutines", gather100)
python bench.py -o before.json # full run; --fast for a quick check
python -m pyperf compare_to before.json after.json --table
Measured: 48.6 ns ± 1.2, 3.15 µs ± 0.06, 1.48 µs ± 0.02 and 150 µs ± 4. The standard deviations are a few percent — what makes a comparison meaningful. compare_to reports whether a difference between two runs is significant, which is the honest way to say "this change made it faster".
Verify: the standard deviation is a small fraction of the mean; if not, the machine is too noisy (step 5).
3. Benchmark the right unit¶
A microbenchmark answers a narrow question. Make it the question your design depends on:
# Question: is a task per item affordable for 10,000 items per request?
async def per_item_tasks(items):
await asyncio.gather(*(asyncio.create_task(process(i)) for i in items))
# Question: what does the extra await layer of a wrapper cost?
async def wrapped(x):
return await inner(x)
The measured costs answer common design questions directly: creating and awaiting a task costs about 3 µs, so a task per item for 10,000 items adds about 30 ms of overhead — significant on a fast path, irrelevant next to network calls; an extra coroutine layer costs about 50 ns, so thin async wrappers are free in practice; yielding with sleep(0) costs about 1.5 µs, so yielding every iteration in a tight CPU loop of a million iterations adds 1.5 s. Benchmark the realistic shape — including argument sizes and the work done — rather than an empty function, once you know the overhead floor.
Verify: the benchmark's input matches the production case it informs (number of items, payload size).
4. Compare alternatives under the same conditions¶
Microbenchmarks are most useful as comparisons — two implementations of the same thing, measured the same way, ideally in the same pyperf run:
async def with_gather(items):
return await asyncio.gather(*(process(i) for i in items))
async def with_taskgroup(items):
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(process(i)) for i in items]
return [t.result() for t in tasks]
ITEMS = list(range(100))
runner.bench_async_func("gather x100", with_gather, ITEMS)
runner.bench_async_func("TaskGroup x100", with_taskgroup, ITEMS)
Pass arguments through bench_async_func rather than constructing data inside the timed coroutine, so setup is not measured. Comparing within one run on one machine removes most environmental variation; comparing numbers from different days or machines does not. Results like these inform choices such as migrating from gather to TaskGroup, where the question is whether structured concurrency costs anything measurable.
Verify: the two variants' results come from the same run and machine, and the difference is larger than their combined standard deviations.
5. Control the environment¶
Small timings are sensitive to CPU frequency scaling, other processes and memory layout. pyperf can help stabilize the machine and detect noise:
python -m pyperf system show # reports frequency scaling, turbo, isolated CPUs
sudo python -m pyperf system tune # pins frequency, disables turbo (dedicated benchmark hosts)
python bench.py --rigorous # more processes and values
python bench.py --affinity=3 # pin worker processes to a quiet CPU
On a shared development machine, run comparisons back to back, prefer --rigorous over --fast for decisions, and repeat if the standard deviation is high. In CI, use the same runner type for baseline and candidate and treat small differences as noise. Remember the limits of microbenchmarks: they say nothing about contention, I/O or memory pressure under load, which need the end-to-end tests in finding the saturation point of an async service.
Verify: pyperf system show reports no warnings on the benchmark host, or the results note that they came from an untuned machine.
Verification¶
Async microbenchmarks are trustworthy when:
- The event loop is created outside the timed region, with many iterations inside.
- pyperf (or repeated runs) reports spread, and it is small relative to the mean.
- Alternatives are compared in the same run, with significance checked.
- Results are applied at realistic scale, and load behaviour is tested separately.
Diagnostic Hook: when a microbenchmark result surprises you, run it again with ten times the iterations per measurement. If the per-iteration time drops, fixed setup costs were inside the timed region — the error that made asyncio.run per call report 44.0 µs for a 3.15 µs operation.
Pitfalls & edge cases¶
asyncio.runper measurement. Measured: 44.0 µs reported for 3.15 µs.- Single runs. Without spread, a 5% difference means nothing.
- Empty functions. They measure overhead, not the realistic workload.
- Comparing across machines or days. Environmental variation swamps small effects.
Frequently Asked Questions¶
How do I benchmark an async function in Python?
Use pyperf's Runner.bench_async_func, which runs the coroutine many times inside one event loop across several processes and reports mean and standard deviation, or time many iterations inside one loop with timeit.
How expensive is asyncio.create_task?
Creating and awaiting a task cost 3.15 µs ± 0.06 in testing with pyperf on Python 3.14; awaiting a coroutine directly cost 48.6 ns.
Why does my asyncio benchmark report such high times?
It probably includes event loop creation: calling asyncio.run per measurement reported 44.0 µs for an operation that costs 3.15 µs.
What does asyncio.sleep(0) cost?
About 1.48 µs per call in testing, so yielding on every iteration of a million-iteration loop adds roughly 1.5 s.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Measuring latency percentiles without averaging them — the statistics side of reporting results.
- Resilience, Cancellation & Error Handling — the section overview.