Using Threads for GIL-Releasing C Extensions¶
"CPU-bound work goes to a process pool" is the standard advice for asyncio services, and for pure-Python code it is right. But much heavy work in real services is not pure Python: hashing uploads, compressing payloads, decoding images, matrix maths. Those run inside C extensions that release the GIL while they work, so plain threads run them in parallel — with no pickling, no extra processes and no shared-memory plumbing. Measured on Python 3.14 with eight threads, SHA-256 over 8 MB buffers ran 5.5× faster than serially, zlib compression 4.6×, a NumPy 600×600 matrix multiply 2.8×, and a pure-Python loop 1.0× — no gain at all. The loop-lag difference was starker: while four threads hashed, the event loop's heartbeat stayed at 0.07 ms median lag; while four threads ran Python loops, it rose to 5.2 ms median and 300 ms p99. This guide tells the two cases apart and routes each correctly.
Prerequisites¶
- Python 3.11+ on a standard (GIL) build; examples use
pip install numpy. - Thread offloading, from running blocking SDK calls with asyncio.to_thread.
- The GIL's effect on services, from how the GIL affects async services.
1. Find out whether your work releases the GIL¶
Do not guess from documentation; measure serial against threaded on your actual operation:
import asyncio
import time
from concurrent.futures import ThreadPoolExecutor
async def speedup(fn, n: int = 8) -> float:
loop = asyncio.get_running_loop()
t = time.perf_counter()
for i in range(n):
fn(i)
serial = time.perf_counter() - t
with ThreadPoolExecutor(max_workers=8) as pool:
t = time.perf_counter()
await asyncio.gather(*(loop.run_in_executor(pool, fn, i) for i in range(n)))
threaded = time.perf_counter() - t
return serial / threaded
Results for four common operations, eight runs each on eight threads:
| Operation | Serial | 8 threads | Speed-up |
|---|---|---|---|
hashlib.sha256 over 8 MB |
0.026 s | 0.005 s | 5.5× |
zlib.compress of ~6 MB JSON |
0.086 s | 0.019 s | 4.6× |
NumPy A @ A, 600×600, BLAS single-threaded |
0.051 s | 0.018 s | 2.8× |
pure-Python for loop |
0.232 s | 0.228 s | 1.0× |
A speed-up close to the thread count means the operation releases the GIL for most of its run time. A speed-up near 1.0 means it holds the GIL, and threads only add switching overhead. Note that NumPy was run with OMP_NUM_THREADS=1 so BLAS did not spawn its own threads; with BLAS multithreading on, a single matrix multiply already uses several cores, and adding Python threads on top oversubscribes the CPU.
Verify: run the harness on your real operation with your real input sizes; the result decides threads or processes.
2. Watch the event loop while threads work¶
Speed is half the story. Threads that hold the GIL compete with the event loop thread for it, and the loop only gets it back at the interpreter's switch interval — 5 ms by default — or when the thread blocks. Threads that release the GIL leave the loop alone:
async def loop_lag_during(fn, threads: int = 4) -> tuple[float, float]:
loop = asyncio.get_running_loop()
lags: list[float] = []
done = False
async def heartbeat():
while not done:
t = time.perf_counter()
await asyncio.sleep(0.001)
lags.append(time.perf_counter() - t - 0.001)
hb = asyncio.create_task(heartbeat())
with ThreadPoolExecutor(threads) as pool:
await asyncio.gather(*(loop.run_in_executor(pool, fn, i) for i in range(threads)))
done = True
await hb
lags.sort()
return lags[len(lags) // 2] * 1000, lags[int(len(lags) * 0.99)] * 1000
Measured: hashing in four threads, loop lag p50 0.07 ms, p99 0.66 ms; Python loops in four threads, p50 5.19 ms, p99 299.6 ms. The second case is the one that surprises people: moving CPU work "off the loop" into a thread made the loop slower for every request, because the GIL was the bottleneck all along. That is exactly why pure-Python CPU work belongs in a process pool, as described in offloading CPU work with run_in_executor.
Verify: run a lag heartbeat during your threaded work; p99 should stay within a few milliseconds.
3. Route each operation to the right executor¶
Make the decision explicit in code, so the next person does not "optimise" a hashing call into a process pool or a parser into a thread:
import hashlib
import zlib
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
gil_free = ThreadPoolExecutor(max_workers=8, thread_name_prefix="gilfree") # C work only
python_cpu = ProcessPoolExecutor(max_workers=8) # Python bytecode
async def digest(data: bytes) -> str:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(gil_free, lambda: hashlib.sha256(data).hexdigest())
async def compress(data: bytes) -> bytes:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(gil_free, zlib.compress, data, 6)
async def score_documents(docs: list[str]) -> list[float]:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(python_cpu, score_all, docs) # pure Python
Threads accept lambdas and closures — nothing is pickled — and arguments are passed by reference, so a 100 MB buffer costs nothing to hand over. That is the other big advantage over processes, where the same buffer would be copied, as measured in sharing large arrays with shared memory.
Verify: grep for each executor's call sites; the thread pool should only wrap calls into C extensions.
4. Know which operations release the GIL¶
Release depends on the extension and sometimes on input size. Common cases in CPython and widely used libraries:
- Usually release:
hashlibdigests for larger inputs,zlib/bz2/lzmacompression, file and socket I/O, most NumPy array operations on large arrays, many image codecs,cryptographyprimitives. - Usually hold: anything that loops in Python,
jsonencoding and decoding with the standard library, regular expressions, pickling, most data-structure manipulation.
Small inputs often do not release at all — releasing and reacquiring the GIL has a cost, so some extensions only release above a size threshold. Measure at production sizes. And remember the inverse risk: an extension that holds the GIL for a long single call blocks the loop thread even when called from a thread, for as long as that call takes.
Verify: for each operation you route to threads, the lag heartbeat stays flat at production input sizes.
5. Revisit the choice on free-threaded Python¶
On free-threaded builds (python3.14t), there is no GIL, and pure-Python threads also run in parallel. The routing rule then changes: threads become viable for Python CPU work too, and the process pool's pickling cost may no longer be worth paying. The evaluation is in evaluating free-threaded Python for CPU-bound threads. Keep the decision in one place — a small module that picks the executor per operation — so it can change with the runtime:
import sys
FREE_THREADED = hasattr(sys, "_is_gil_enabled") and not sys._is_gil_enabled()
python_cpu = (ThreadPoolExecutor(max_workers=8) if FREE_THREADED
else ProcessPoolExecutor(max_workers=8))
Verify: the same service runs correctly under both builds, with the executor chosen at startup and logged.
Verification¶
Executor routing is right when:
- Every thread-pool operation shows a real speed-up with threads, measured at production sizes.
- Loop lag stays flat while thread-pool work runs.
- Pure-Python CPU work runs in processes on GIL builds.
- BLAS or OpenMP threads are not oversubscribed by Python threads on top.
Diagnostic Hook: export loop-lag p99 alongside the number of busy threads in each executor. A lag that rises with busy threads in the "GIL-free" pool means something holding the GIL has been routed there — usually a pure-Python pre- or post-processing step wrapped around the C call. Profile that pool's threads with py-spy to find it.
Pitfalls & edge cases¶
- Python pre-processing around the C call. A loop that builds the input holds the GIL; only the C call itself releases it.
- BLAS oversubscription. NumPy with multithreaded BLAS plus Python threads runs more threads than cores.
- Small inputs. Many extensions do not release the GIL for tiny inputs; batch them.
- Assuming all C code releases the GIL. Many extensions never do; measure.
Frequently Asked Questions¶
Can threads speed up CPU-bound work in Python?
Yes, when the work runs in C code that releases the GIL, such as hashlib, zlib or large NumPy operations: eight threads gave 2.8 to 5.5 times speed-ups in testing. Pure-Python loops gained nothing.
Does hashlib release the GIL?
Yes for larger inputs. SHA-256 over 8 MB buffers ran 5.5 times faster on eight threads than serially, and the event loop's lag stayed under a millisecond while it ran.
Why did moving CPU work to a thread make my asyncio service slower?
If the work is pure Python, the thread holds the GIL and the event loop thread can only get it back at the switch interval. In testing, loop lag rose to 5 ms median and 300 ms p99. Use a process pool for such work.
Should I use threads or processes for NumPy?
Threads, for operations on large arrays that release the GIL, because nothing is copied. Make sure BLAS is not also multithreaded, or limit it with OMP_NUM_THREADS, to avoid oversubscribing cores.
Related¶
- CPU-Bound Task Offloading — up to the topic overview.
- Sizing the default thread pool executor — sizing the thread pool these calls use.
- Concurrent Execution & Worker Patterns — the section overview.