Measuring Thread Pool Executor Saturation in asyncio¶
asyncio.to_thread and run_in_executor hand blocking work to a thread pool with a fixed number of threads. When more work arrives than threads, calls queue — and none of the usual metrics show it, because the call's total time looks like slow work. Measured on Python 3.14 on a 24-core machine, where the default executor has 28 threads, with each call sleeping 50 ms: 20 concurrent calls waited 0.1 ms in the queue at the median; 100 calls waited 49.9 ms at the median and 149.8 ms at p99; 300 calls waited 251.7 ms and 501.6 ms — while every call still ran for 50.1 ms once started. The default executor is also where asyncio performs DNS resolution: loop.getaddrinfo("localhost") took 1.01 ms with an idle executor and 1,951 ms with 60 one-second jobs queued ahead of it. Moving those jobs to their own executor brought DNS back to 0.23 ms. This guide instruments queue wait, exports saturation, and separates executors.
Prerequisites¶
- Python 3.11+ asyncio.
- Executor sizing, from sizing the default thread pool executor.
- The topic overview, Observability & Tracing.
1. Split queue wait from run time¶
Record when work is submitted, when a thread starts it, and when it finishes. The gap between the first two is queueing:
def _timed(fn, args, submitted: float, name: str):
started = time.perf_counter()
executor_wait.labels(name).observe(started - submitted)
try:
return fn(*args)
finally:
executor_run.labels(name).observe(time.perf_counter() - started)
async def run_blocking(fn, *args, name: str = "default", executor=None):
loop = asyncio.get_running_loop()
return await loop.run_in_executor(executor, _timed, fn, args, time.perf_counter(), name)
Measured with 50 ms blocking calls on the 28-thread default executor: run time was 50.1 ms at every concurrency, but queue wait went from 0.1 ms at 20 concurrent calls, to 49.9 ms median and 149.8 ms p99 at 100, to 251.7 ms median and 501.6 ms p99 at 300. Without the split, a dashboard would show the calls "getting slower" and invite optimizing code that was not slow. The same separation for asyncio queues is described in measuring queue wait and service time separately.
Verify: dashboards show executor queue wait and run time as separate series, per executor.
2. Export a saturation gauge¶
Queue wait tells you saturation happened; a gauge shows it as it happens. Count work in flight against the executor's size:
class InstrumentedExecutor(concurrent.futures.ThreadPoolExecutor):
def __init__(self, max_workers: int, name: str):
super().__init__(max_workers, thread_name_prefix=name)
self.name, self.in_flight = name, 0
self._lock = threading.Lock()
def submit(self, fn, /, *args, **kwargs):
with self._lock:
self.in_flight += 1
future = super().submit(fn, *args, **kwargs)
future.add_done_callback(self._done)
return future
def _done(self, _):
with self._lock:
self.in_flight -= 1
def saturation(self) -> float:
return self.in_flight / self._max_workers # > 1.0 means work is queueing
A value above 1.0 is queued work: at 300 concurrent 50 ms calls on 28 threads, about 10.7 at the start of the burst. Sample it every few seconds into a gauge and alert when it stays above 1 — sustained queueing means the executor is undersized or something is blocking in it far longer than expected.
Verify: under a load test, the gauge rises above 1.0 exactly when queue-wait percentiles rise.
3. Watch for DNS behind application work¶
asyncio resolves hostnames with getaddrinfo in the loop's default executor. Every open_connection, create_connection and many HTTP client connections pass through it — the same executor that asyncio.to_thread uses:
t = time.perf_counter()
await loop.getaddrinfo("localhost", 80)
print((time.perf_counter() - t) * 1000) # 1.01 ms idle; 1,951 ms behind 60 x 1 s jobs
Measured: with 60 one-second blocking jobs submitted to the default executor just before the lookup, resolving localhost took 1,951 ms instead of 1.01 ms — it waited for two waves of 28 jobs to finish. In a service, slow file I/O or a blocking SDK call in to_thread shows up as connection timeouts to unrelated hosts. DNS resolution and its caching are covered in resolving DNS without blocking the executor.
Verify: connection setup latency is measured while the default executor is saturated, and stays flat.
4. Give heavy work its own executor¶
Keep the default executor for asyncio's own needs — DNS and short calls — and give each kind of heavy blocking work its own, sized and instrumented separately:
FILES = InstrumentedExecutor(8, "files")
SDK = InstrumentedExecutor(16, "sdk")
async def read_report(path):
return await run_blocking(Path(path).read_bytes, name="files", executor=FILES)
async def put_object(key, body):
return await run_blocking(s3.put_object, key, body, name="sdk", executor=SDK)
Measured: with the same 60 one-second jobs on a separate 16-thread executor, getaddrinfo took 0.23 ms. Separate executors are bulkheads for threads: a burst of slow SDK calls can saturate its own pool without delaying file reads or connection setup elsewhere. Size each from its own measured run time and arrival rate, and shut them down in the application's lifespan.
Verify: a saturation test of one executor leaves queue wait in the others, and DNS latency, unchanged.
5. Size from wait, not from guesses¶
With queue wait and run time measured per executor, sizing becomes arithmetic: an executor needs about arrival rate × run time threads to avoid queueing, plus headroom for bursts — the same Little's-law relationship used in rate-limited worker pools:
def threads_needed(calls_per_second: float, run_time_s: float, headroom: float = 1.5) -> int:
return math.ceil(calls_per_second * run_time_s * headroom)
threads_needed(400, 0.05) # 400 calls/s of 50 ms each -> 30 threads
More threads cost memory and, for CPU-bound Python, add GIL contention rather than throughput, so a rising run time alongside rising thread count is the sign to move the work to processes instead. Re-check after changes: a dependency that gets slower raises run time, and the same arrival rate then saturates the pool.
Verify: each executor's size is derived from its measured arrival rate and run time, and saturation stays below 1.0 at normal peak.
Verification¶
Executor saturation is visible when:
- Queue wait and run time are measured separately for every executor.
- A saturation gauge shows in-flight work against capacity.
- Heavy blocking work runs on its own executors, leaving the default one for DNS.
- Executor sizes follow from measured arrival rates and run times.
Diagnostic Hook: when outbound connections to healthy services start timing out while a service is doing file or SDK work in threads, time loop.getaddrinfo during the incident. Resolution taking seconds instead of a millisecond — 1,951 ms against 1.01 ms here — means DNS is queued behind application work on the default executor.
Pitfalls & edge cases¶
- Timing only the whole
to_threadcall. Measured: 50 ms of work looked like up to 550 ms. - Sharing the default executor with slow work. Measured: DNS delayed to 1,951 ms.
- Adding threads for CPU-bound Python. Run time rises with the GIL; use processes.
- Reading
_work_queue.qsize()in production. It is private; count in-flight work instead.
Frequently Asked Questions¶
How many threads does asyncio's default executor have?
min(32, os.cpu_count() + 4): 28 on this 24-core machine. Beyond that, to_thread calls queue: 300 concurrent 50 ms calls waited 252 ms at the median.
How do I measure asyncio.to_thread queue time?
Record time.perf_counter() when submitting, again when the function starts in the thread, and the difference is queue wait; record run time separately.
Why are my DNS lookups slow in asyncio?
getaddrinfo runs in the default executor. With 60 one-second jobs queued there, resolving localhost took 1,951 ms instead of 1.01 ms.
Should I use a separate ThreadPoolExecutor for blocking work?
For heavy or slow work, yes: with the jobs on their own executor, DNS took 0.23 ms during the same load.
Related¶
- Observability & Tracing — up to the topic overview.
- Reporting background task errors — another failure that hides from default metrics.
- Resilience, Cancellation & Error Handling — the section overview.