How the GIL Affects Async Services¶
An asyncio service runs all its handlers on one thread, and the GIL means only one thread runs Python bytecode at a time. Together these decide how CPU work in a request affects every other request. Measured in a service handling a stream of light requests (a 5 ms await each) alongside heavier requests that parsed a JSON document for 22 ms: with the parse inline in the handler, the light requests' p99 latency rose from the ideal 5 ms to 55 ms; moving the parse to a thread with asyncio.to_thread only improved it to 47 ms, because the thread still needed the GIL; moving it to a process pool brought p99 to 6.1 ms. The median stayed at 5.1 ms in all three cases — the damage is entirely in the tail, which is why it is often missed. This guide explains the mechanism and how to keep CPU work from leaking into everyone's latency.
Prerequisites¶
- Python 3.11+ on a standard (GIL) build, stdlib only.
- Event loop lag, from measuring event loop lag in production.
- Process pools from asyncio, from offloading CPU work with run_in_executor.
1. See why inline CPU work hurts the tail¶
The event loop runs one callback at a time. A handler that parses for 22 ms holds the loop for 22 ms; every other request whose timer or socket became ready during that time waits until the parse finishes:
async def handle_import(request):
body = await request.body()
rows = json.loads(body) # 22 ms on the loop thread: nothing else runs
await store(rows)
return Response(status_code=202)
The light requests' median did not move — most of them never coincided with a parse — but p99 rose elevenfold to 55 ms, the latency of a request that arrived just as one or two parses started. In production this shows up as a p99 that tracks the arrival rate of one expensive endpoint, while that endpoint's own latency looks fine.
Verify: run a loop-lag heartbeat; spikes equal to the parse duration confirm the cause.
2. Understand why a thread barely helps¶
Moving the parse to asyncio.to_thread frees the loop in principle: the handler awaits a future while another thread parses. But json.loads from the standard library runs Python-level C code that holds the GIL for its whole duration, and the loop thread needs the GIL to run anything. The interpreter forces a switch between threads that want the GIL every 5 ms (sys.getswitchinterval()), so the loop gets slices of time rather than freedom:
import sys
print(sys.getswitchinterval()) # 0.005
Measured, the thread version's p99 was 46.9 ms — barely better than inline. A long-running C call that holds the GIL without checking the switch interval can block the loop thread for the call's entire duration, so for some workloads moving to a thread does nothing at all. Contrast this with work that releases the GIL — hashing, compression, NumPy — where threads do help, as measured in using threads for GIL-releasing C extensions.
Verify: compare loop lag with the work inline and in a thread; if both are similar, the work holds the GIL and needs a process.
3. Move GIL-holding work to processes¶
A process has its own interpreter and its own GIL, so its CPU work cannot contend with the loop thread:
from concurrent.futures import ProcessPoolExecutor
parse_pool = ProcessPoolExecutor(max_workers=2)
def parse_and_validate(raw: bytes) -> list[dict]: # runs in the worker process
rows = json.loads(raw)
return [validate(r) for r in rows]
async def handle_import(request):
body = await request.body()
loop = asyncio.get_running_loop()
rows = await loop.run_in_executor(parse_pool, parse_and_validate, body)
await store(rows)
return Response(status_code=202)
Measured: light-request p99 6.1 ms. The cost is moving data between processes — the body is pickled to the worker and the result pickled back — so offload the whole CPU-heavy step, parse and validate and transform, in one call rather than several small ones. For large inputs that cost can dominate; see reducing pickle overhead in ProcessPoolExecutor payloads.
Verify: under load, loop lag stays near zero while the heavy endpoint is busy, and the pool's queue does not grow without bound.
4. Size the service around CPU per request¶
Even with offloading, each request costs some CPU on the loop thread — routing, serialisation, logging. That per-request CPU sets a hard ceiling per process: at 1 ms of loop-thread CPU per request, one process cannot exceed about 1,000 requests per second, however concurrent the code is. Measure it and plan processes accordingly:
import time
class CpuPerRequest:
def __init__(self) -> None:
self.last_cpu = time.process_time()
self.last_count = 0
def sample(self, completed: int) -> float:
cpu = time.process_time()
per = (cpu - self.last_cpu) / max(1, completed - self.last_count)
self.last_cpu, self.last_count = cpu, completed
return per * 1000 # ms of CPU per request
process_time() counts CPU across all threads in the process, so offloaded thread work shows up here too; process-pool work does not. Divide your target throughput by the per-process ceiling to get the number of worker processes, and keep each process's CPU utilisation well below 100% to leave headroom for the event loop to stay responsive. This is the calculation behind sizing uvicorn workers for async services.
Verify: at peak, per-process CPU stays below about 70% and loop lag stays low.
5. Re-evaluate on free-threaded builds¶
Python 3.13 introduced experimental free-threaded builds (python3.13t, python3.14t) without the GIL. On those, a pure-Python parse in a thread runs in parallel with the loop thread, and to_thread becomes an effective offload for CPU work without pickling. Two caveats: single-threaded code runs somewhat slower on free-threaded builds, and every C extension you depend on must support them. Measure the same p99 experiment there before changing the design:
import sys
gil_enabled = getattr(sys, "_is_gil_enabled", lambda: True)()
cpu_executor = None if not gil_enabled else ProcessPoolExecutor(max_workers=2)
# None -> loop.run_in_executor uses the default ThreadPoolExecutor
The decision logic is the same as before — keep heavy work off the loop thread — only the cheapest correct place to put it changes. The broader evaluation is in evaluating free-threaded Python for CPU-bound threads.
Verify: on a free-threaded build, the thread variant's p99 approaches the process variant's; if it does not, an extension is re-enabling the GIL.
Verification¶
CPU work is isolated correctly when:
- Light requests' p99 stays near their ideal while heavy endpoints are busy.
- Loop lag does not track the arrival rate of any single endpoint.
- GIL-holding work runs in processes, GIL-releasing work in threads.
- Per-request loop-thread CPU is measured and used to size worker processes.
Diagnostic Hook: export loop lag p99 alongside per-endpoint request rates. A correlation between lag and one endpoint's rate identifies the endpoint doing CPU work on the loop — or in a thread that holds the GIL — without profiling. Confirm with a py-spy sample filtered to the loop thread.
Pitfalls & edge cases¶
- Watching only the median. It did not move at all while p99 rose elevenfold.
to_threadfor pure-Python CPU work. It barely helps under the GIL.- Many small offloads. Each process hop pickles data; offload whole steps.
- Assuming free-threading. Check
sys._is_gil_enabled()at runtime; an extension can turn the GIL back on.
Frequently Asked Questions¶
Does the GIL matter for asyncio services?
Yes, whenever requests do CPU work. All handlers share the event loop thread, and threads you offload to share the GIL with it. In testing, a 22 ms parse raised other requests' p99 from 5 ms to 55 ms inline and 47 ms in a thread.
Why didn't asyncio.to_thread fix my latency?
The work holds the GIL, so the event loop thread can only run in slices at the 5 ms switch interval while the worker thread computes. Use a process pool for Python-level CPU work.
How much CPU per request can an asyncio process handle?
Roughly one second of loop-thread CPU per second, so at 1 ms per request a process tops out near 1,000 requests per second. Keep utilisation well below that to keep latency stable.
Does free-threaded Python change this?
On free-threaded builds, threads run Python code in parallel, so to_thread can offload CPU work without pickling. Measure first: single-threaded code is somewhat slower and extensions must support it.
Related¶
- Threading vs Multiprocessing vs Asyncio — up to the topic overview.
- Combining asyncio with multiprocessing for mixed workloads — the architecture this leads to.
- Concurrent Execution & Worker Patterns — the section overview.