When asyncio Is Slower Than Threads¶
asyncio's advantage is concurrency over slow waits: thousands of network calls that each spend milliseconds idle. Its costs are paid per await, per task and per thread hop, and when each operation is fast those costs dominate. The clearest example is SQLite: measured on Python 3.14, a simple indexed lookup took 8.4 µs through the standard sqlite3 module, 30.4 µs when each query was sent to a dedicated thread with run_in_executor, and 58.8 µs through aiosqlite, which runs every operation on a background thread and hops back for each call — seven times slower for the same query. Even pure coroutine overhead shows up: a three-level call chain cost 55 ns as plain functions and 149 ns as coroutines. This guide identifies the workloads where async code is slower and what to do instead.
Prerequisites¶
- Python 3.11+; the SQLite comparison uses
pip install aiosqlite. - The model trade-offs, from Threading vs Multiprocessing vs Asyncio.
- Thread hop costs, from running blocking SDK calls with asyncio.to_thread.
1. Recognise the pattern: fast operations behind a thread hop¶
Many "async" libraries for local or embedded resources are thread wrappers: every call is shipped to a worker thread and the result shipped back to the loop. When the operation itself takes microseconds, the round trip is the cost:
import asyncio
import sqlite3
import time
import aiosqlite # pip install aiosqlite
N = 10_000
def sync_queries(path: str) -> float:
conn = sqlite3.connect(path)
t = time.perf_counter()
for i in range(N):
conn.execute("select v from t where id = ?", (i % 1000 + 1,)).fetchone()
return (time.perf_counter() - t) / N
async def aiosqlite_queries(path: str) -> float:
async with aiosqlite.connect(path) as conn:
t = time.perf_counter()
for i in range(N):
async with conn.execute("select v from t where id = ?", (i % 1000 + 1,)) as cur:
await cur.fetchone()
return (time.perf_counter() - t) / N
Results: 8.4 µs per query synchronously, 58.8 µs with aiosqlite — which pays a hop for execute and another for fetchone. A bare run_in_executor per query cost 30.4 µs: a single hop is roughly 20 µs on this machine, and that is the floor for any thread-backed async API. For a database on the same machine answering in microseconds, the async wrapper adds latency and gains nothing, because there is no slow wait to overlap.
Verify: compare per-operation latency of the async wrapper with the sync library it wraps, at your real query mix.
2. Know the per-await and per-task floor¶
Coroutines are cheap, not free. Each await of a coroutine adds frame setup and the coroutine protocol on top of a normal call:
async def leaf(x): return x + 1
async def mid(x): return await leaf(x)
async def top(x): return await mid(x)
def leaf_s(x): return x + 1
def mid_s(x): return leaf_s(x)
def top_s(x): return mid_s(x)
A million calls: 55 ns per sync chain, 149 ns per async chain — about 30 ns extra per level. Creating a task costs far more, around 2 µs including scheduling, as measured in scheduling callbacks with call_soon vs create_task. None of this matters for code that awaits network I/O. It matters for hot inner loops: a parser written as async def all the way down, or a pipeline that creates a task per small item, spends its time on the machinery.
Keep pure computation synchronous. Make a function async only if it awaits something.
Verify: profile a hot path with py-spy; frames inside asyncio and coroutine machinery should be a small fraction of samples.
3. Prefer threads when the work is a few blocking calls¶
Threads beat asyncio when concurrency is low and the libraries are blocking:
- A CLI or batch job making a handful of concurrent calls with a sync SDK. A
ThreadPoolExecutorwith ten threads is simpler than an async rewrite, and ten threads cost nothing worth measuring. - Code dominated by a blocking library with no async equivalent. Wrapping every call in
to_threadadds a hop to every call; running the whole worker in a thread pool removes them. - Short-lived scripts. Starting an event loop, an async client and its pool for three requests is overhead with no payback.
from concurrent.futures import ThreadPoolExecutor
import boto3
s3 = boto3.client("s3")
def download_all(keys: list[str]) -> None:
with ThreadPoolExecutor(max_workers=16) as pool:
list(pool.map(lambda k: s3.download_file(BUCKET, k, f"/tmp/{k}"), keys))
That is the right shape for a batch job; the async alternative and its trade-offs are in wrapping sync cloud SDKs with asyncio.to_thread.
Verify: if every await in a piece of async code is a to_thread call, the threaded version is simpler and at least as fast.
4. Watch for CPU work hiding in async handlers¶
asyncio runs everything on one thread. A handler that spends 5 ms parsing JSON or rendering a template blocks every other request on that loop for 5 ms. Threads contend for the GIL too, but the interpreter switches between them every few milliseconds; the event loop never preempts a running handler. In an asyncio service, CPU work per request bounds throughput at roughly 1 / cpu_time_per_request per process, regardless of concurrency:
async def handler(request):
body = await request.body()
data = json.loads(body) # 5 ms for a large body: the loop is blocked
result = transform(data) # pure Python: more blocking
return JSONResponse(result)
At 5 ms of CPU per request, one process tops out near 200 requests per second, and every concurrent request's latency includes the CPU time of the ones ahead of it. The fixes are the usual ones: more processes, a process pool for heavy transforms, and keeping handlers I/O-bound — described in how the GIL affects async services.
Verify: measure event loop lag under load; if it tracks request rate, CPU work in handlers is the limit.
5. Decide with a measurement, not a reputation¶
When choosing a model for a component, measure the operation the component will do most, at realistic concurrency:
async def compare(op_async, op_sync, n: int = 1_000, concurrency: int = 50):
sem = asyncio.Semaphore(concurrency)
async def one():
async with sem:
await op_async()
t = time.perf_counter()
await asyncio.gather(*(one() for _ in range(n)))
async_time = time.perf_counter() - t
t = time.perf_counter()
with ThreadPoolExecutor(concurrency) as pool:
list(pool.map(lambda _: op_sync(), range(n)))
return async_time, time.perf_counter() - t
Run it for the operations that matter — the database query, the outbound call, the cache read — and let the numbers decide per component. The broader decision framework, including when processes are the answer, is in choosing a concurrency model for a WebSocket gateway.
Verify: each component's concurrency model is backed by a recorded measurement of its dominant operation.
Verification¶
The concurrency model fits when:
- Async components spend most of their time awaiting slow I/O, verified by profiling.
- Thread-backed async wrappers are not used for microsecond operations in hot paths.
- Pure computation is synchronous, not
async defwithout awaits. - Choices are backed by measurements of each component's dominant operation.
Diagnostic Hook: profile production with py-spy and compute the share of samples in asyncio internals, to_thread/executor plumbing and coroutine machinery. More than a few percent means the service is paying for async overhead it is not using; a high share in executor plumbing specifically points at a thread-backed async library sitting on a hot path.
Pitfalls & edge cases¶
- "Async everything." Converting pure functions to coroutines adds overhead and buys nothing.
- Async wrappers over local, fast resources. SQLite, local files on SSD, in-process caches.
- Task per tiny item. Thousands of microsecond tasks spend more time scheduling than working.
- Ignoring CPU per request. It caps throughput per process no matter how concurrent the code is.
Frequently Asked Questions¶
Is asyncio always faster than threads?
No. asyncio is faster when many slow network waits can overlap. For fast local operations it adds overhead: an indexed SQLite lookup took 8.4 µs with sqlite3 and 58.8 µs with aiosqlite in testing.
Why is aiosqlite slower than sqlite3?
It runs each operation on a background thread and hands the result back to the event loop, so every call pays thread hops that cost more than the query itself on a local database.
How much overhead does async/await add?
About 30 ns per level of coroutine call on Python 3.14 — a three-level chain took 149 ns against 55 ns for plain functions — and around 2 µs per task created. It is negligible next to network I/O and significant in tight loops.
When should I use threads instead of asyncio?
For a small number of concurrent blocking calls, for code dominated by blocking libraries without async versions, and for short scripts. Use asyncio when you need thousands of concurrent slow I/O operations.
Related¶
- Threading vs Multiprocessing vs Asyncio — up to the topic overview.
- Using SQLite from asyncio with aiosqlite — getting reasonable performance when you must.
- Concurrent Execution & Worker Patterns — the section overview.