Sizing asyncio's Default Thread Pool Executor¶
Every asyncio.to_thread() call and every loop.run_in_executor(None, ...) call shares one thread pool: the loop's default executor, created lazily on first use with min(32, os.process_cpu_count() + 4) threads. On a 24-core machine that is 28 threads. A hundred concurrent to_thread calls that each block for 100 ms therefore do not take 100 ms — in a test they took 0.40 s, four waves of 28. Setting a 100-thread default executor brought the same batch down to 0.11 s. The default is sized for CPU-ish work, and most code that ends up in to_thread is blocking I/O: a sync SDK, a DNS lookup, a file read. This guide shows how to see the queueing, choose a size, and when one shared pool is the wrong design.
Prerequisites¶
- Python 3.11+, stdlib only; numbers below are from Python 3.14 on a 24-core Linux machine.
- What
to_threaddoes, from running blocking SDK calls with asyncio.to_thread. - Loop setup, from Event Loop Configuration.
1. Find out what you have¶
The size is computed when the executor is first created, from the CPUs available to the process. On Python 3.13+ that uses os.process_cpu_count(), which respects CPU affinity; container CPU quotas are not affinity, so a container limited to 2 CPUs on a 64-core host still gets min(32, 64 + 4) = 32 threads.
import asyncio
import os
async def main() -> None:
loop = asyncio.get_running_loop()
await asyncio.to_thread(lambda: None) # force lazy creation
executor = loop._default_executor # private, but stable for inspection
print("cpus:", os.process_cpu_count(), "max_workers:", executor._max_workers)
asyncio.run(main())
# cpus: 24 max_workers: 28
Note the attribute names are private; use them for a one-off check, not in production code. The point is to know the number, because every blocking call in the process competes for it — your own to_thread calls, libraries that use run_in_executor(None, ...) internally, and on most platforms loop.getaddrinfo() for DNS resolution.
Verify: print the size in your deployment environment, not on your laptop; containers and CI runners often differ.
2. Measure queue wait, not just call time¶
A saturated pool does not fail; calls just take longer, and the extra time looks like the blocking call being slow. Separate the two by timing from submission to start and from start to finish:
import asyncio
import time
from contextvars import copy_context
async def timed_to_thread(fn, *args, metrics, name: str):
submitted = time.perf_counter()
def wrapper():
started = time.perf_counter()
metrics.observe("executor_queue_wait_seconds", started - submitted, call=name)
try:
return fn(*args)
finally:
metrics.observe("executor_run_seconds", time.perf_counter() - started, call=name)
loop = asyncio.get_running_loop()
ctx = copy_context()
return await loop.run_in_executor(None, ctx.run, wrapper)
Queue wait near zero means the pool is big enough. Queue wait that grows with load while run time stays flat is the signature of an undersized pool — the same distinction made for request queues in measuring queue wait and service time separately.
Verify: run 100 concurrent 100 ms blocking calls; with the default pool the p99 queue wait should be close to 300 ms (three waves ahead of the last call).
3. Size the pool for the blocking you actually do¶
For blocking I/O, the threads mostly wait, and the right size comes from Little's law: concurrent blocking calls = arrival rate × time each call blocks. At 200 calls per second blocking 50 ms each, you need about 10 threads busy on average; give it headroom for bursts and slow periods, say 3–4× the average.
Set it once, before anything uses the pool:
import asyncio
from concurrent.futures import ThreadPoolExecutor
async def main() -> None:
loop = asyncio.get_running_loop()
loop.set_default_executor(
ThreadPoolExecutor(max_workers=64, thread_name_prefix="blocking-io")
)
await serve()
asyncio.run(main())
The thread name prefix pays for itself the first time you read a py-spy dump or a thread list and need to know which threads belong to this pool, as in profiling asyncio applications with py-spy.
Do not set it to hundreds "to be safe". Each thread reserves a stack (8 MB of virtual memory by default on Linux, a fraction of that resident), and every thread contends for the GIL whenever it does Python work. If the blocking calls are actually CPU-bound Python, more threads make things slower, and the work belongs in a process pool instead.
Verify: after resizing, queue wait under peak load stays near zero, and threading.active_count() stays below max_workers plus your other threads.
4. Split pools by workload¶
One shared pool means one slow dependency can starve every other blocking call. If a blocking SDK starts taking 30 s per call during an outage, it occupies every thread, and DNS lookups for healthy dependencies queue behind it. Give slow or risky work its own executor:
from concurrent.futures import ThreadPoolExecutor
s3_pool = ThreadPoolExecutor(max_workers=16, thread_name_prefix="s3")
files_pool = ThreadPoolExecutor(max_workers=8, thread_name_prefix="files")
async def upload(path: str, key: str) -> None:
loop = asyncio.get_running_loop()
await loop.run_in_executor(s3_pool, s3_client.upload_file, path, BUCKET, key)
async def read_config(path: str) -> bytes:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(files_pool, _read, path)
This is the bulkhead pattern applied to threads: a stalled S3 client can exhaust its own 16 threads and nothing else. Keep asyncio.to_thread and the default executor for short, well-behaved calls.
Verify: make one dependency hang; calls through the other pools and through the default executor keep completing at normal latency.
5. Shut the pool down with the loop¶
asyncio.run() calls loop.shutdown_default_executor() on exit, which waits for running calls to finish — since Python 3.12 with a timeout of 300 s, after which it warns and moves on. Executors you created yourself are not shut down for you:
async def main() -> None:
try:
await serve()
finally:
loop = asyncio.get_running_loop()
await loop.run_in_executor(None, s3_pool.shutdown, True) # wait for in-flight
files_pool.shutdown(wait=False, cancel_futures=True) # drop queued work
cancel_futures=True drops calls that have not started; calls already running in a thread cannot be interrupted and run to completion. A blocking call that never returns therefore holds shutdown open — the case covered in timing out blocking calls in threads.
Verify: stop the service during load; it exits within your shutdown budget and logs no "executor did not finish joining its threads" warning.
Verification¶
The executor is sized correctly when:
- Queue wait stays near zero at peak load, measured per call site.
- Threads in use stay below
max_workersmost of the time, with headroom for bursts. - A hung dependency only exhausts its own pool, not the default executor.
- Shutdown completes within its budget with no thread-join warnings.
Diagnostic Hook: export executor queue wait as a histogram and the number of busy threads as a gauge per pool. Alert when p99 queue wait exceeds 10% of the calls' own run time, or when a pool sits at 100% busy for more than a minute — that pool is either undersized or wrapping a dependency that has stalled.
Pitfalls & edge cases¶
- Setting the executor after first use.
set_default_executorreplaces the pool for future calls; the old pool and its threads stay alive until shut down. - Reading
os.cpu_count()in a container. It reports the host's CPUs, not your quota; size explicitly. - CPU work in the default pool. It contends for the GIL with the event loop thread and slows everything; use processes.
- DNS starvation.
getaddrinfouses the default executor; a saturated pool makes new connections slow for reasons that look like network problems — see caching DNS lookups in async HTTP clients. - Calling
to_threadfrom many loops. Each loop has its own default executor; two loops in one process mean two pools.
Frequently Asked Questions¶
How many threads does asyncio.to_thread use?
It uses the loop's default ThreadPoolExecutor, which has min(32, CPU count + 4) threads unless you replace it — 28 on a 24-core machine. Calls beyond that queue until a thread is free.
How do I increase the asyncio default executor size?
Create a ThreadPoolExecutor with the max_workers you want and pass it to loop.set_default_executor at startup, before anything calls to_thread or run_in_executor with None.
Why are my to_thread calls slow under load?
Usually because the default pool is saturated and calls are queueing for a thread. Measure the time from submission to start separately from the run time; if the queue wait grows with load, the pool is too small or one slow dependency is occupying it.
Should I use one thread pool or several?
Use the default pool for short, reliable calls and give slow or failure-prone dependencies their own executor, so one stalled dependency cannot starve every other blocking call in the process.
Related¶
- Event Loop Configuration — up to the topic overview.
- Optimizing worker pool sizes for mixed I/O and CPU workloads — the sizing arithmetic in more depth.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.