Avoiding Deadlocks When Threads Wait on the Loop¶
Code that bridges threads and asyncio waits in two directions: threads block on run_coroutine_threadsafe(...).result() for the loop, and coroutines await to_thread or executor futures for threads. Each direction is safe on its own. Combine them carelessly and you get a deadlock that shows up only under specific timing, which in production means "the service hangs once a week". Two shapes account for almost all of them, and both were reproduced with timeouts as the only way out: calling .result() on a run_coroutine_threadsafe future from the loop thread itself waited the full 1-second timeout, because the thread that must run the coroutine was the one blocked waiting for it; and a worker thread that held a threading.Lock while waiting on a coroutine that needed the same lock stalled both the worker and the event loop until the worker's 1-second timeout released it. This guide shows how each arises, how to detect them, and the designs that make them impossible.
Prerequisites¶
- Python 3.11+, stdlib only.
- The background-loop bridge, from running an event loop in a background thread.
- Shared state rules, from how to safely share state between async tasks and threads.
1. Recognise deadlock 1: waiting on the loop from the loop¶
A synchronous helper that submits work to the loop and blocks for the result works perfectly when called from a worker thread. Called — directly or indirectly — from code already running on the loop, it blocks the only thread that could run the coroutine:
import asyncio
import concurrent.futures as cf
def fetch_sync(loop: asyncio.AbstractEventLoop, url: str, timeout: float = 1.0):
"""Sync wrapper used by legacy code. Safe from threads; fatal from the loop thread."""
fut = asyncio.run_coroutine_threadsafe(fetch(url), loop)
return fut.result(timeout)
async def handler(request):
loop = asyncio.get_running_loop()
data = fetch_sync(loop, request.url) # blocks the loop thread waiting for the loop
Reproduced: the call waited the full 1 s and raised TimeoutError. Without a timeout it would hang forever, and so would every other request on that loop. The call path is usually less obvious than this — a sync library callback that happens to be invoked on the loop thread, or a "sync" utility that someone later called from an async view.
Verify: add a guard (step 3) to every sync bridge and run the test suite; any call from the loop thread now fails loudly instead of hanging.
2. Recognise deadlock 2: holding a lock across the bridge¶
The second shape involves a lock. A worker thread takes a threading.Lock, then waits for a coroutine; the coroutine — running on the loop thread — tries to take the same lock:
import threading
state_lock = threading.Lock()
def worker(loop):
with state_lock: # holds the lock...
fut = asyncio.run_coroutine_threadsafe(update_cache(), loop)
return fut.result(timeout=1) # ...while waiting for the loop
async def update_cache():
with state_lock: # blocks the LOOP THREAD on it
cache.refresh()
Reproduced: the worker timed out after 1 s with the lock held, and for that whole second the loop thread sat blocked inside with state_lock, so no other coroutine ran. Once the worker's timeout fired and it released the lock, update_cache proceeded — after its caller had already given up. The general rule is the classic one, applied across the thread/loop boundary: never wait on another execution context while holding a lock that context may need. Note also that update_cache blocks the loop thread on a threading.Lock at all, which is a hazard even without the deadlock — loop-side code should not take locks that threads hold for long.
Verify: search for run_coroutine_threadsafe(...).result() and call_soon_threadsafe calls made inside with <lock> blocks; each is a candidate.
3. Detect calls from the loop thread¶
Make the first deadlock fail fast instead of hanging. A sync bridge can check whether it is being called on the loop it would wait for:
def run_sync(coro, loop: asyncio.AbstractEventLoop, timeout: float = 30.0):
try:
running = asyncio.get_running_loop()
except RuntimeError:
running = None
if running is loop:
coro.close() # avoid a never-awaited warning
raise RuntimeError("run_sync() called from the event loop it waits on; "
"await the coroutine instead")
fut = asyncio.run_coroutine_threadsafe(coro, loop)
try:
return fut.result(timeout)
except cf.TimeoutError:
fut.cancel()
raise
asyncio.get_running_loop() returns the loop only when called from the loop's own thread while it is running, so the check is exact. Raising immediately turns a hang into a stack trace that names the offending caller. Combine it with debug mode in tests, which also flags non-thread-safe calls, as described in enabling asyncio debug mode in tests and CI.
Verify: a test that calls run_sync from inside a coroutine gets the RuntimeError immediately, not a timeout.
4. Design so the waits cannot cycle¶
Detection catches mistakes; design prevents them. Three rules, each of which removes a whole class of deadlock:
- The loop never blocks on threads. Loop-side code awaits
to_threador executor futures; it never takes athreading.Lockthat threads hold for more than microseconds, and never calls a sync bridge. - Threads never hold locks while waiting for the loop. Copy what you need under the lock, release it, then wait.
- State has one owner. Mutations to shared state happen on the loop (threads submit requests) or in threads (the loop submits work), never both under a shared lock.
def worker(loop):
with state_lock:
snapshot = dict(shared_state) # copy under the lock
# lock released before crossing to the loop
fut = asyncio.run_coroutine_threadsafe(update_cache(snapshot), loop)
return fut.result(timeout=5)
async def update_cache(snapshot: dict) -> None:
cache.refresh(snapshot) # loop-owned state, no thread lock needed
With the loop as the sole owner of cache, there is no lock for anyone to hold across the boundary. The ownership pattern is the same one recommended in why asyncio locks are not thread-safe.
Verify: review every lock acquired in code that also crosses the thread/loop boundary; none is held at the crossing.
5. Always put timeouts on cross-boundary waits¶
Even with good design, put a timeout on every .result() and every await of an executor future. A deadlock with a timeout is an error with a traceback; without one, it is a hung process that a health check eventually kills, with no evidence of why:
fut.result(timeout=5) # thread waiting for the loop
async with asyncio.timeout(5):
await loop.run_in_executor(pool, blocking_call) # loop waiting for a thread
Pick timeouts from the operation's expected duration plus margin, not a global default. And when a timeout fires on a bridge call, log both the caller's thread name and, if possible, a stack dump of the loop thread — the techniques in dumping stacks of a hung asyncio program show which side was blocked.
Verify: grep for .result() with no argument on concurrent futures; there should be none in service code.
Verification¶
Bridge code is deadlock-free when:
- Sync bridges refuse to run on their own loop, raising immediately.
- No lock is held while waiting across the boundary.
- Every cross-boundary wait has a timeout, and timeouts cancel the waited-on work.
- State has a single owner, so no lock is needed at the boundary.
Diagnostic Hook: count timeouts on bridge calls per call site, and on every timeout capture a faulthandler dump of all threads. A timeout whose dump shows the loop thread inside a threading.Lock acquire is deadlock 2; one where the caller is the loop thread is deadlock 1. Either way the dump names the line.
Pitfalls & edge cases¶
- Indirect calls. A sync library callback invoked on the loop thread can reach a sync bridge several frames down.
fut.result()with no timeout. Converts any deadlock into a permanent hang.- Taking
threading.Lockin coroutines. Blocks the loop thread whenever a worker holds it. - Reentrant locks.
RLockdoes not help across threads; the loop thread is a different thread.
Frequently Asked Questions¶
Why does run_coroutine_threadsafe(...).result() hang?
Most often because it was called on the event loop's own thread. The coroutine can only run on that thread, which is blocked waiting for it. Await the coroutine directly when you are already in async code.
Can a threading.Lock deadlock asyncio code?
Yes. If a thread holds the lock while waiting for a coroutine that also needs it, the coroutine blocks the event loop thread on the lock and neither side progresses. Release locks before waiting across the boundary.
How do I detect a sync-to-async call made from the event loop thread?
Check whether asyncio.get_running_loop() returns the loop you are about to wait on; if it does, raise immediately instead of blocking.
Should cross-thread asyncio waits have timeouts?
Always. A deadlock with a timeout becomes an exception with a traceback; without one it is a silent hang. Cancel the waited-on work when the timeout fires.
Related¶
- Hybrid Concurrency Models — up to the topic overview.
- Bridging queues between threads and asyncio tasks — message passing that avoids shared locks.
- Concurrent Execution & Worker Patterns — the section overview.