Skip to content

Measuring asyncio Speedups Between Python Versions

Each Python release since 3.11 has made asyncio's own machinery faster, and upgrade notes quote the gains. Whether your service sees any of it depends on how much of its time is spent in that machinery rather than in system calls, the network or your own code. Measured with one benchmark script run under uv-managed CPython 3.10.20, 3.11.15, 3.12.13, 3.13.14 and 3.14.6, best of five runs each: creating and awaiting a task fell from 13.5 µs on 3.10 to 3.2 µs on 3.14, 4.3× faster; gather over 100,000 coroutines from 17.9 to 5.7 µs per coroutine, 3.1×; a Queue put/get pair from 0.89 to 0.41 µs; await asyncio.sleep(0) from 2.32 to 1.52 µs. A TCP echo round trip over loopback, dominated by system calls, improved only from 24.6 to 19.8 µs, 1.24×. Most of the task gain arrived in two steps — 3.11 and 3.13 — with 3.12 slightly slower than 3.11 on several operations. This guide builds a benchmark that answers the question for your own workload, before and after an upgrade.

Prerequisites

1. Write one benchmark that runs on every version

The benchmark must use only APIs available on your oldest interpreter, and each case should isolate one operation your service does a lot of. Time per operation, take the best of several runs, and print machine-readable output:

import asyncio
import json
import sys
import time


async def nop():
    return None


async def tasks(n):
    pending = [asyncio.ensure_future(nop()) for _ in range(n)]
    for t in pending:
        await t


async def gather(n):
    await asyncio.gather(*(nop() for _ in range(n)))


async def queue(n):
    q = asyncio.Queue(100)

    async def produce():
        for i in range(n):
            await q.put(i)

    async def consume():
        for _ in range(n):
            await q.get()

    await asyncio.gather(produce(), consume())


CASES = [("create+await task", tasks, 100_000),
         ("gather", gather, 100_000),
         ("Queue put/get", queue, 100_000)]

results = {}
for name, fn, n in CASES:
    best = float("inf")
    for _ in range(5):
        start = time.perf_counter()
        asyncio.run(fn(n))
        best = min(best, (time.perf_counter() - start) / n * 1e6)
    results[name] = round(best, 2)

print(json.dumps({"python": sys.version.split()[0], "us_per_op": results}))

TaskGroup and asyncio.timeout are deliberately absent: they do not exist on 3.10, and a benchmark that cannot run on the old version cannot measure the upgrade. Taking the minimum of five runs discards most interference from other processes — the measurements here were taken on a shared 24-core machine with a load average between 15 and 19.

Verify: the script runs unchanged under uv run --python 3.10 and under your newest target.

Create and await one task, by Python version 6 horizontal bars comparing 3.10 with the others. Create and await one task, by Python version 3.10 13.5 us 3.11 8.2 us 3.12 8.5 us 3.13 5.1 us 3.14 3.2 us 3.14t 2.4 us Best of five runs of 100,000 tasks; uv-managed CPython builds on one machine. 4.3x faster from 3.10 to 3.14, arriving mostly in 3.11 and 3.13.

2. Run it under each interpreter

With uv, each interpreter is one command and needs no virtual environment for a standard-library benchmark:

for v in 3.10 3.11 3.12 3.13 3.14; do
    uv run --no-project --python "$v" bench.py
done > bench.jsonl

Measured on this machine, the whole sweep — six cases, five runs each, five versions — took about 75 seconds. Run every version in the same session, one after another, so that machine load affects them alike; a 3.10 result from Monday and a 3.14 result from Friday compare two machines, not two interpreters. For a service with dependencies, create a virtual environment per version with uv venv --python X.Y and install the same lock file in each, so you compare interpreters rather than library versions.

Verify: bench.jsonl has one line per interpreter and each line names the exact patch version.

3. Read where the gains are

Lay the results side by side and compute the ratio between your current and target versions:

import json

rows = [json.loads(line) for line in open("bench.jsonl")]
base = rows[0]["us_per_op"]
for row in rows:
    ratios = {k: f"{base[k] / v:.2f}x" for k, v in row["us_per_op"].items()}
    print(row["python"], ratios)

Measured, from 3.10 to 3.14: create and await a task 4.3×, gather 3.1×, Queue put/get 2.2×, uncontended Lock acquire and release 1.8× (0.37 to 0.21 µs), sleep(0) 1.5×, TCP echo round trip 1.24×. The pattern is consistent: the more an operation is pure asyncio bookkeeping, the more it gained; the more it is a system call, the less. Between neighbouring versions the picture is uneven — 3.12 measured slightly slower than 3.11 for tasks (8.5 against 8.2 µs) and gather (10.9 against 10.6 µs) — so an upgrade of one minor version can show no gain at all, and a skipped version can contain all of it. The free-threaded 3.14t build, run with the same script, measured 2.4 µs per task and 5.0 µs per gather item, at or slightly better than the default build.

Verify: for each case you know the ratio between your current and target version.

Microseconds per operation, best of five A grid of 6 rows by 6 columns. Microseconds per operation, best of five operation 3.10 3.11 3.12 3.13 3.14 await sleep(0) 2.32 1.86 1.71 1.60 1.52 create+await task 13.5 8.2 8.5 5.1 3.2 gather, per coroutine 17.9 10.6 10.9 6.3 5.7 Queue put/get 0.89 0.52 0.54 0.47 0.41 uncontended Lock 0.37 0.28 0.28 0.25 0.21 TCP echo round trip 24.6 20.1 20.8 20.3 19.8 Bookkeeping got much faster; system calls did not.

4. Translate microbenchmarks into your service's time

A 4× faster task matters only in proportion to how much of a request's time is spent creating tasks. Estimate the share from a profile, or count the operations per request and multiply:

# per request, from a profile or by counting
tasks_per_request = 20
queue_ops_per_request = 40
round_trips_per_request = 3

def asyncio_cost_us(t):
    return (tasks_per_request * t["create+await task"]
            + queue_ops_per_request * t["Queue put/get"]
            + round_trips_per_request * t["TCP echo round trip"])

With the measured figures, this hypothetical request spends 380 µs on these operations under 3.10 and 139 µs under 3.14. If the request takes 20 ms end to end — mostly waiting on a database — the upgrade saves about 1.2% of its latency; if it takes 1 ms of mostly in-process work, it saves about 24%. A fan-out service that creates thousands of short tasks per request gains far more than a proxy that spends its time in send and recv. The only reliable number is a load test of the real service on both interpreters, as in load testing async services with Locust, but the arithmetic tells you whether the load test is worth running.

Verify: you have an estimate of the share of request time spent in asyncio operations, and a load test confirms it for at least one endpoint.

How much will this service gain from a newer Python? A decision on Where does a request spend its time with 4 outcomes. How much will this service gain from a newer Python? Where does a request spend its time? creating many short tasks large gain tasks 4.3x faster passing items through queues moderate gain Queue 2.2x faster send and recv on sockets small gain round trip 1.24x faster waiting on a database or upstream little change the wait dominates Measured from 3.10.20 to 3.14.6 on the same machine.

5. Keep the benchmark as an upgrade gate

Store the benchmark in the repository and run it when a new interpreter or patch release is evaluated. The comparison that matters before an upgrade is not "is the new version faster in general" but "is any operation we depend on slower":

import json
import sys

old, new = (json.loads(line)["us_per_op"] for line in open(sys.argv[1]))
regressions = {k: f"{new[k] / old[k]:.2f}x slower"
               for k in old if new[k] > old[k] * 1.10}
print(regressions or "no operation more than 10% slower")
sys.exit(1 if regressions else 0)

Applied to 3.11 against 3.12, this gate flagged nothing at the 10% threshold — the largest step backwards was task creation at 1.04× slower — and applied to 3.11 against 3.13 every case improved except the TCP round trip, which was 1% slower — within noise. Ten per cent is a deliberate margin: on a shared machine, best-of-five figures still moved by a few per cent between sweeps, and a tighter threshold would flag noise. Pair the gate with the functional checks in testing asyncio code across Python versions, so an upgrade is judged on both behaviour and speed.

Verify: the gate runs against the target interpreter before an upgrade is merged.

Verification

You know what an upgrade will do to your service's asyncio performance when:

  • One benchmark runs unchanged on every interpreter you compare.
  • All versions are measured in one session, best of several runs.
  • Microbenchmark ratios are translated into request time using per-request operation counts.
  • A regression gate checks the target version before the upgrade lands.

Diagnostic Hook: after an interpreter upgrade, compare event-loop CPU time per request — process CPU time divided by requests served — before and after the deploy. A drop confirms the gain; no change means request time is dominated by waits and system calls, exactly as the arithmetic in step 4 predicts.

Pitfalls & edge cases

  • Expecting every version to be faster. 3.12 measured slightly slower than 3.11 for tasks.
  • Benchmarks that use new APIs. They cannot run on the version you are leaving.
  • Comparing runs from different days. Machine load moves results more than some upgrades.
  • Reading a 4× microbenchmark as a 4× service. A TCP round trip gained only 1.24×.

Frequently Asked Questions

Is asyncio faster in Python 3.14?

Its bookkeeping is: task creation was 4.3x faster than on 3.10 and 1.6x faster than on 3.13 in testing, and gather 3.1x faster than on 3.10. Operations dominated by system calls, such as socket round trips, improved much less.

Which Python version gave the biggest asyncio speedup?

In this benchmark, 3.11 and 3.13. Task creation went from 13.5 to 8.2 µs at 3.11 and from 8.5 to 5.1 µs at 3.13, while 3.12 was marginally slower than 3.11.

Will upgrading Python make my async web service faster?

Only in proportion to the time it spends in asyncio operations. A request dominated by database waits gained about 1% in the worked estimate; one dominated by in-process task work gained about 24%.

How do I benchmark asyncio across Python versions?

Write one script using only APIs from your oldest version, run it with uv run --python X.Y for each version in the same session, take the best of several runs, and compare per-operation ratios.