Skip to content

How the GIL Affects Async Services

An asyncio service runs all its handlers on one thread, and the GIL means only one thread runs Python bytecode at a time. Together these decide how CPU work in a request affects every other request. Measured in a service handling a stream of light requests (a 5 ms await each) alongside heavier requests that parsed a JSON document for 22 ms: with the parse inline in the handler, the light requests' p99 latency rose from the ideal 5 ms to 55 ms; moving the parse to a thread with asyncio.to_thread only improved it to 47 ms, because the thread still needed the GIL; moving it to a process pool brought p99 to 6.1 ms. The median stayed at 5.1 ms in all three cases — the damage is entirely in the tail, which is why it is often missed. This guide explains the mechanism and how to keep CPU work from leaking into everyone's latency.

Prerequisites

1. See why inline CPU work hurts the tail

The event loop runs one callback at a time. A handler that parses for 22 ms holds the loop for 22 ms; every other request whose timer or socket became ready during that time waits until the parse finishes:

async def handle_import(request):
    body = await request.body()
    rows = json.loads(body)                  # 22 ms on the loop thread: nothing else runs
    await store(rows)
    return Response(status_code=202)

The light requests' median did not move — most of them never coincided with a parse — but p99 rose elevenfold to 55 ms, the latency of a request that arrived just as one or two parses started. In production this shows up as a p99 that tracks the arrival rate of one expensive endpoint, while that endpoint's own latency looks fine.

Verify: run a loop-lag heartbeat; spikes equal to the parse duration confirm the cause.

p99 of light requests while 22 ms parses run 3 horizontal bars comparing parse inline on the loop with the others. p99 of light requests while 22 ms parses run parse inline on the loop 55.3 ms parse in a thread (to_thread) 46.9 ms parse in a process pool 6.1 ms Median stayed at 5.1 ms in every case; Python 3.14 standard build. A thread barely helps pure-Python CPU work, because the loop thread still waits for the GIL.

2. Understand why a thread barely helps

Moving the parse to asyncio.to_thread frees the loop in principle: the handler awaits a future while another thread parses. But json.loads from the standard library runs Python-level C code that holds the GIL for its whole duration, and the loop thread needs the GIL to run anything. The interpreter forces a switch between threads that want the GIL every 5 ms (sys.getswitchinterval()), so the loop gets slices of time rather than freedom:

import sys
print(sys.getswitchinterval())     # 0.005

Measured, the thread version's p99 was 46.9 ms — barely better than inline. A long-running C call that holds the GIL without checking the switch interval can block the loop thread for the call's entire duration, so for some workloads moving to a thread does nothing at all. Contrast this with work that releases the GIL — hashing, compression, NumPy — where threads do help, as measured in using threads for GIL-releasing C extensions.

Verify: compare loop lag with the work inline and in a thread; if both are similar, the work holds the GIL and needs a process.

3. Move GIL-holding work to processes

A process has its own interpreter and its own GIL, so its CPU work cannot contend with the loop thread:

from concurrent.futures import ProcessPoolExecutor

parse_pool = ProcessPoolExecutor(max_workers=2)


def parse_and_validate(raw: bytes) -> list[dict]:     # runs in the worker process
    rows = json.loads(raw)
    return [validate(r) for r in rows]


async def handle_import(request):
    body = await request.body()
    loop = asyncio.get_running_loop()
    rows = await loop.run_in_executor(parse_pool, parse_and_validate, body)
    await store(rows)
    return Response(status_code=202)

Measured: light-request p99 6.1 ms. The cost is moving data between processes — the body is pickled to the worker and the result pickled back — so offload the whole CPU-heavy step, parse and validate and transform, in one call rather than several small ones. For large inputs that cost can dominate; see reducing pickle overhead in ProcessPoolExecutor payloads.

Verify: under load, loop lag stays near zero while the heavy endpoint is busy, and the pool's queue does not grow without bound.

Where the GIL sits in each design A sequence of 6 messages between 4 participants. Where the GIL sits in each design loop thread GIL worker thread worker process holds GIL: json.loads needs GIL: waits up to 5 ms switch interval: brief slice back to the parse parses under its own GIL runs freely while the process works Threads share one GIL with the loop; a process brings its own.

4. Size the service around CPU per request

Even with offloading, each request costs some CPU on the loop thread — routing, serialisation, logging. That per-request CPU sets a hard ceiling per process: at 1 ms of loop-thread CPU per request, one process cannot exceed about 1,000 requests per second, however concurrent the code is. Measure it and plan processes accordingly:

import time


class CpuPerRequest:
    def __init__(self) -> None:
        self.last_cpu = time.process_time()
        self.last_count = 0

    def sample(self, completed: int) -> float:
        cpu = time.process_time()
        per = (cpu - self.last_cpu) / max(1, completed - self.last_count)
        self.last_cpu, self.last_count = cpu, completed
        return per * 1000                                 # ms of CPU per request

process_time() counts CPU across all threads in the process, so offloaded thread work shows up here too; process-pool work does not. Divide your target throughput by the per-process ceiling to get the number of worker processes, and keep each process's CPU utilisation well below 100% to leave headroom for the event loop to stay responsive. This is the calculation behind sizing uvicorn workers for async services.

Verify: at peak, per-process CPU stays below about 70% and loop lag stays low.

5. Re-evaluate on free-threaded builds

Python 3.13 introduced experimental free-threaded builds (python3.13t, python3.14t) without the GIL. On those, a pure-Python parse in a thread runs in parallel with the loop thread, and to_thread becomes an effective offload for CPU work without pickling. Two caveats: single-threaded code runs somewhat slower on free-threaded builds, and every C extension you depend on must support them. Measure the same p99 experiment there before changing the design:

import sys

gil_enabled = getattr(sys, "_is_gil_enabled", lambda: True)()
cpu_executor = None if not gil_enabled else ProcessPoolExecutor(max_workers=2)
# None -> loop.run_in_executor uses the default ThreadPoolExecutor

The decision logic is the same as before — keep heavy work off the loop thread — only the cheapest correct place to put it changes. The broader evaluation is in evaluating free-threaded Python for CPU-bound threads.

Verify: on a free-threaded build, the thread variant's p99 approaches the process variant's; if it does not, an extension is re-enabling the GIL.

Where should this CPU work run? A decision on How long, and does it release the GIL with 3 outcomes. Where should this CPU work run? How long, and does it release the GIL? well under 1 ms inline not worth a hop releases the GIL to_thread no pickling holds the GIL, ms or more process pool own interpreter The median hides this problem; decide by p99 and loop lag.

Verification

CPU work is isolated correctly when:

  • Light requests' p99 stays near their ideal while heavy endpoints are busy.
  • Loop lag does not track the arrival rate of any single endpoint.
  • GIL-holding work runs in processes, GIL-releasing work in threads.
  • Per-request loop-thread CPU is measured and used to size worker processes.

Diagnostic Hook: export loop lag p99 alongside per-endpoint request rates. A correlation between lag and one endpoint's rate identifies the endpoint doing CPU work on the loop — or in a thread that holds the GIL — without profiling. Confirm with a py-spy sample filtered to the loop thread.

Pitfalls & edge cases

  • Watching only the median. It did not move at all while p99 rose elevenfold.
  • to_thread for pure-Python CPU work. It barely helps under the GIL.
  • Many small offloads. Each process hop pickles data; offload whole steps.
  • Assuming free-threading. Check sys._is_gil_enabled() at runtime; an extension can turn the GIL back on.

Frequently Asked Questions

Does the GIL matter for asyncio services?

Yes, whenever requests do CPU work. All handlers share the event loop thread, and threads you offload to share the GIL with it. In testing, a 22 ms parse raised other requests' p99 from 5 ms to 55 ms inline and 47 ms in a thread.

Why didn't asyncio.to_thread fix my latency?

The work holds the GIL, so the event loop thread can only run in slices at the 5 ms switch interval while the worker thread computes. Use a process pool for Python-level CPU work.

How much CPU per request can an asyncio process handle?

Roughly one second of loop-thread CPU per second, so at 1 ms per request a process tops out near 1,000 requests per second. Keep utilisation well below that to keep latency stable.

Does free-threaded Python change this?

On free-threaded builds, threads run Python code in parallel, so to_thread can offload CPU work without pickling. Measure first: single-threaded code is somewhat slower and extensions must support it.