Measuring asyncio Speedups Between Python Versions¶
Each Python release since 3.11 has made asyncio's own machinery faster, and upgrade notes quote the gains. Whether your service sees any of it depends on how much of its time is spent in that machinery rather than in system calls, the network or your own code. Measured with one benchmark script run under uv-managed CPython 3.10.20, 3.11.15, 3.12.13, 3.13.14 and 3.14.6, best of five runs each: creating and awaiting a task fell from 13.5 µs on 3.10 to 3.2 µs on 3.14, 4.3× faster; gather over 100,000 coroutines from 17.9 to 5.7 µs per coroutine, 3.1×; a Queue put/get pair from 0.89 to 0.41 µs; await asyncio.sleep(0) from 2.32 to 1.52 µs. A TCP echo round trip over loopback, dominated by system calls, improved only from 24.6 to 19.8 µs, 1.24×. Most of the task gain arrived in two steps — 3.11 and 3.13 — with 3.12 slightly slower than 3.11 on several operations. This guide builds a benchmark that answers the question for your own workload, before and after an upgrade.
Prerequisites¶
- uv, to install and run several interpreters with
uv run --python X.Y. - Benchmarking hygiene, from microbenchmarking coroutines with pyperf and timeit.
- The topic overview, Asyncio Across Python Versions.
1. Write one benchmark that runs on every version¶
The benchmark must use only APIs available on your oldest interpreter, and each case should isolate one operation your service does a lot of. Time per operation, take the best of several runs, and print machine-readable output:
import asyncio
import json
import sys
import time
async def nop():
return None
async def tasks(n):
pending = [asyncio.ensure_future(nop()) for _ in range(n)]
for t in pending:
await t
async def gather(n):
await asyncio.gather(*(nop() for _ in range(n)))
async def queue(n):
q = asyncio.Queue(100)
async def produce():
for i in range(n):
await q.put(i)
async def consume():
for _ in range(n):
await q.get()
await asyncio.gather(produce(), consume())
CASES = [("create+await task", tasks, 100_000),
("gather", gather, 100_000),
("Queue put/get", queue, 100_000)]
results = {}
for name, fn, n in CASES:
best = float("inf")
for _ in range(5):
start = time.perf_counter()
asyncio.run(fn(n))
best = min(best, (time.perf_counter() - start) / n * 1e6)
results[name] = round(best, 2)
print(json.dumps({"python": sys.version.split()[0], "us_per_op": results}))
TaskGroup and asyncio.timeout are deliberately absent: they do not exist on 3.10, and a benchmark that cannot run on the old version cannot measure the upgrade. Taking the minimum of five runs discards most interference from other processes — the measurements here were taken on a shared 24-core machine with a load average between 15 and 19.
Verify: the script runs unchanged under uv run --python 3.10 and under your newest target.
2. Run it under each interpreter¶
With uv, each interpreter is one command and needs no virtual environment for a standard-library benchmark:
for v in 3.10 3.11 3.12 3.13 3.14; do
uv run --no-project --python "$v" bench.py
done > bench.jsonl
Measured on this machine, the whole sweep — six cases, five runs each, five versions — took about 75 seconds. Run every version in the same session, one after another, so that machine load affects them alike; a 3.10 result from Monday and a 3.14 result from Friday compare two machines, not two interpreters. For a service with dependencies, create a virtual environment per version with uv venv --python X.Y and install the same lock file in each, so you compare interpreters rather than library versions.
Verify: bench.jsonl has one line per interpreter and each line names the exact patch version.
3. Read where the gains are¶
Lay the results side by side and compute the ratio between your current and target versions:
import json
rows = [json.loads(line) for line in open("bench.jsonl")]
base = rows[0]["us_per_op"]
for row in rows:
ratios = {k: f"{base[k] / v:.2f}x" for k, v in row["us_per_op"].items()}
print(row["python"], ratios)
Measured, from 3.10 to 3.14: create and await a task 4.3×, gather 3.1×, Queue put/get 2.2×, uncontended Lock acquire and release 1.8× (0.37 to 0.21 µs), sleep(0) 1.5×, TCP echo round trip 1.24×. The pattern is consistent: the more an operation is pure asyncio bookkeeping, the more it gained; the more it is a system call, the less. Between neighbouring versions the picture is uneven — 3.12 measured slightly slower than 3.11 for tasks (8.5 against 8.2 µs) and gather (10.9 against 10.6 µs) — so an upgrade of one minor version can show no gain at all, and a skipped version can contain all of it. The free-threaded 3.14t build, run with the same script, measured 2.4 µs per task and 5.0 µs per gather item, at or slightly better than the default build.
Verify: for each case you know the ratio between your current and target version.
4. Translate microbenchmarks into your service's time¶
A 4× faster task matters only in proportion to how much of a request's time is spent creating tasks. Estimate the share from a profile, or count the operations per request and multiply:
# per request, from a profile or by counting
tasks_per_request = 20
queue_ops_per_request = 40
round_trips_per_request = 3
def asyncio_cost_us(t):
return (tasks_per_request * t["create+await task"]
+ queue_ops_per_request * t["Queue put/get"]
+ round_trips_per_request * t["TCP echo round trip"])
With the measured figures, this hypothetical request spends 380 µs on these operations under 3.10 and 139 µs under 3.14. If the request takes 20 ms end to end — mostly waiting on a database — the upgrade saves about 1.2% of its latency; if it takes 1 ms of mostly in-process work, it saves about 24%. A fan-out service that creates thousands of short tasks per request gains far more than a proxy that spends its time in send and recv. The only reliable number is a load test of the real service on both interpreters, as in load testing async services with Locust, but the arithmetic tells you whether the load test is worth running.
Verify: you have an estimate of the share of request time spent in asyncio operations, and a load test confirms it for at least one endpoint.
5. Keep the benchmark as an upgrade gate¶
Store the benchmark in the repository and run it when a new interpreter or patch release is evaluated. The comparison that matters before an upgrade is not "is the new version faster in general" but "is any operation we depend on slower":
import json
import sys
old, new = (json.loads(line)["us_per_op"] for line in open(sys.argv[1]))
regressions = {k: f"{new[k] / old[k]:.2f}x slower"
for k in old if new[k] > old[k] * 1.10}
print(regressions or "no operation more than 10% slower")
sys.exit(1 if regressions else 0)
Applied to 3.11 against 3.12, this gate flagged nothing at the 10% threshold — the largest step backwards was task creation at 1.04× slower — and applied to 3.11 against 3.13 every case improved except the TCP round trip, which was 1% slower — within noise. Ten per cent is a deliberate margin: on a shared machine, best-of-five figures still moved by a few per cent between sweeps, and a tighter threshold would flag noise. Pair the gate with the functional checks in testing asyncio code across Python versions, so an upgrade is judged on both behaviour and speed.
Verify: the gate runs against the target interpreter before an upgrade is merged.
Verification¶
You know what an upgrade will do to your service's asyncio performance when:
- One benchmark runs unchanged on every interpreter you compare.
- All versions are measured in one session, best of several runs.
- Microbenchmark ratios are translated into request time using per-request operation counts.
- A regression gate checks the target version before the upgrade lands.
Diagnostic Hook: after an interpreter upgrade, compare event-loop CPU time per request — process CPU time divided by requests served — before and after the deploy. A drop confirms the gain; no change means request time is dominated by waits and system calls, exactly as the arithmetic in step 4 predicts.
Pitfalls & edge cases¶
- Expecting every version to be faster. 3.12 measured slightly slower than 3.11 for tasks.
- Benchmarks that use new APIs. They cannot run on the version you are leaving.
- Comparing runs from different days. Machine load moves results more than some upgrades.
- Reading a 4× microbenchmark as a 4× service. A TCP round trip gained only 1.24×.
Frequently Asked Questions¶
Is asyncio faster in Python 3.14?
Its bookkeeping is: task creation was 4.3x faster than on 3.10 and 1.6x faster than on 3.13 in testing, and gather 3.1x faster than on 3.10. Operations dominated by system calls, such as socket round trips, improved much less.
Which Python version gave the biggest asyncio speedup?
In this benchmark, 3.11 and 3.13. Task creation went from 13.5 to 8.2 µs at 3.11 and from 8.5 to 5.1 µs at 3.13, while 3.12 was marginally slower than 3.11.
Will upgrading Python make my async web service faster?
Only in proportion to the time it spends in asyncio operations. A request dominated by database waits gained about 1% in the worked estimate; one dominated by in-process task work gained about 24%.
How do I benchmark asyncio across Python versions?
Write one script using only APIs from your oldest version, run it with uv run --python X.Y for each version in the same session, take the best of several runs, and compare per-operation ratios.
Related¶
- Asyncio Across Python Versions — up to the topic overview.
- Supporting several Python versions in async libraries — shipping code that runs on all of them.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.