Skip to content

Avoiding Deadlocks When Threads Wait on the Loop

Code that bridges threads and asyncio waits in two directions: threads block on run_coroutine_threadsafe(...).result() for the loop, and coroutines await to_thread or executor futures for threads. Each direction is safe on its own. Combine them carelessly and you get a deadlock that shows up only under specific timing, which in production means "the service hangs once a week". Two shapes account for almost all of them, and both were reproduced with timeouts as the only way out: calling .result() on a run_coroutine_threadsafe future from the loop thread itself waited the full 1-second timeout, because the thread that must run the coroutine was the one blocked waiting for it; and a worker thread that held a threading.Lock while waiting on a coroutine that needed the same lock stalled both the worker and the event loop until the worker's 1-second timeout released it. This guide shows how each arises, how to detect them, and the designs that make them impossible.

Prerequisites

1. Recognise deadlock 1: waiting on the loop from the loop

A synchronous helper that submits work to the loop and blocks for the result works perfectly when called from a worker thread. Called — directly or indirectly — from code already running on the loop, it blocks the only thread that could run the coroutine:

import asyncio
import concurrent.futures as cf


def fetch_sync(loop: asyncio.AbstractEventLoop, url: str, timeout: float = 1.0):
    """Sync wrapper used by legacy code. Safe from threads; fatal from the loop thread."""
    fut = asyncio.run_coroutine_threadsafe(fetch(url), loop)
    return fut.result(timeout)


async def handler(request):
    loop = asyncio.get_running_loop()
    data = fetch_sync(loop, request.url)       # blocks the loop thread waiting for the loop

Reproduced: the call waited the full 1 s and raised TimeoutError. Without a timeout it would hang forever, and so would every other request on that loop. The call path is usually less obvious than this — a sync library callback that happens to be invoked on the loop thread, or a "sync" utility that someone later called from an async view.

Verify: add a guard (step 3) to every sync bridge and run the test suite; any call from the loop thread now fails loudly instead of hanging.

Deadlock 1 - the loop thread waits for itself A sequence of 5 messages between 3 participants. Deadlock 1 - the loop thread waits for itself handler on loop thread sync wrapper event loop fetch_sync(loop, url) run_coroutine_threadsafe(fetch()) fut.result(): blocks the loop thread fetch() queued, but the loop thread is blocked TimeoutError after 1 s, or hang forever The coroutine can only run on the thread that is waiting for it.

2. Recognise deadlock 2: holding a lock across the bridge

The second shape involves a lock. A worker thread takes a threading.Lock, then waits for a coroutine; the coroutine — running on the loop thread — tries to take the same lock:

import threading

state_lock = threading.Lock()


def worker(loop):
    with state_lock:                                            # holds the lock...
        fut = asyncio.run_coroutine_threadsafe(update_cache(), loop)
        return fut.result(timeout=1)                            # ...while waiting for the loop


async def update_cache():
    with state_lock:                                            # blocks the LOOP THREAD on it
        cache.refresh()

Reproduced: the worker timed out after 1 s with the lock held, and for that whole second the loop thread sat blocked inside with state_lock, so no other coroutine ran. Once the worker's timeout fired and it released the lock, update_cache proceeded — after its caller had already given up. The general rule is the classic one, applied across the thread/loop boundary: never wait on another execution context while holding a lock that context may need. Note also that update_cache blocks the loop thread on a threading.Lock at all, which is a hazard even without the deadlock — loop-side code should not take locks that threads hold for long.

Verify: search for run_coroutine_threadsafe(...).result() and call_soon_threadsafe calls made inside with <lock> blocks; each is a candidate.

Deadlock 2 - a lock held across the thread-to-loop wait 3 lanes over time. Deadlock 2 - a lock held across the thread-to-loop wait worker thread take lock wait for coroutine, lock held timeout, release loop thread start coroutine blocked on the lock proceeds, too late other coroutines none can run time → Reproduced: both sides stalled for the full 1 s timeout, and the whole event loop with them.

3. Detect calls from the loop thread

Make the first deadlock fail fast instead of hanging. A sync bridge can check whether it is being called on the loop it would wait for:

def run_sync(coro, loop: asyncio.AbstractEventLoop, timeout: float = 30.0):
    try:
        running = asyncio.get_running_loop()
    except RuntimeError:
        running = None
    if running is loop:
        coro.close()                                   # avoid a never-awaited warning
        raise RuntimeError("run_sync() called from the event loop it waits on; "
                           "await the coroutine instead")
    fut = asyncio.run_coroutine_threadsafe(coro, loop)
    try:
        return fut.result(timeout)
    except cf.TimeoutError:
        fut.cancel()
        raise

asyncio.get_running_loop() returns the loop only when called from the loop's own thread while it is running, so the check is exact. Raising immediately turns a hang into a stack trace that names the offending caller. Combine it with debug mode in tests, which also flags non-thread-safe calls, as described in enabling asyncio debug mode in tests and CI.

Verify: a test that calls run_sync from inside a coroutine gets the RuntimeError immediately, not a timeout.

4. Design so the waits cannot cycle

Detection catches mistakes; design prevents them. Three rules, each of which removes a whole class of deadlock:

  • The loop never blocks on threads. Loop-side code awaits to_thread or executor futures; it never takes a threading.Lock that threads hold for more than microseconds, and never calls a sync bridge.
  • Threads never hold locks while waiting for the loop. Copy what you need under the lock, release it, then wait.
  • State has one owner. Mutations to shared state happen on the loop (threads submit requests) or in threads (the loop submits work), never both under a shared lock.
def worker(loop):
    with state_lock:
        snapshot = dict(shared_state)                 # copy under the lock
    # lock released before crossing to the loop
    fut = asyncio.run_coroutine_threadsafe(update_cache(snapshot), loop)
    return fut.result(timeout=5)


async def update_cache(snapshot: dict) -> None:
    cache.refresh(snapshot)                           # loop-owned state, no thread lock needed

With the loop as the sole owner of cache, there is no lock for anyone to hold across the boundary. The ownership pattern is the same one recommended in why asyncio locks are not thread-safe.

Verify: review every lock acquired in code that also crosses the thread/loop boundary; none is held at the crossing.

Is this cross-boundary wait safe? A decision on Where is the waiting code, and what does it hold with 3 outcomes. Is this cross-boundary wait safe? Where is the waiting code, and what does it hold? on the loop thread, waiting on it deadlock 1 await instead a thread holding a shared lock deadlock 2 copy, release, then wait a thread, nothing held safe with a timeout Every bridge wait should pass this check; timeouts turn the failures into errors rather than hangs.

5. Always put timeouts on cross-boundary waits

Even with good design, put a timeout on every .result() and every await of an executor future. A deadlock with a timeout is an error with a traceback; without one, it is a hung process that a health check eventually kills, with no evidence of why:

fut.result(timeout=5)                                  # thread waiting for the loop

async with asyncio.timeout(5):
    await loop.run_in_executor(pool, blocking_call)    # loop waiting for a thread

Pick timeouts from the operation's expected duration plus margin, not a global default. And when a timeout fires on a bridge call, log both the caller's thread name and, if possible, a stack dump of the loop thread — the techniques in dumping stacks of a hung asyncio program show which side was blocked.

Verify: grep for .result() with no argument on concurrent futures; there should be none in service code.

Verification

Bridge code is deadlock-free when:

  • Sync bridges refuse to run on their own loop, raising immediately.
  • No lock is held while waiting across the boundary.
  • Every cross-boundary wait has a timeout, and timeouts cancel the waited-on work.
  • State has a single owner, so no lock is needed at the boundary.

Diagnostic Hook: count timeouts on bridge calls per call site, and on every timeout capture a faulthandler dump of all threads. A timeout whose dump shows the loop thread inside a threading.Lock acquire is deadlock 2; one where the caller is the loop thread is deadlock 1. Either way the dump names the line.

Pitfalls & edge cases

  • Indirect calls. A sync library callback invoked on the loop thread can reach a sync bridge several frames down.
  • fut.result() with no timeout. Converts any deadlock into a permanent hang.
  • Taking threading.Lock in coroutines. Blocks the loop thread whenever a worker holds it.
  • Reentrant locks. RLock does not help across threads; the loop thread is a different thread.

Frequently Asked Questions

Why does run_coroutine_threadsafe(...).result() hang?

Most often because it was called on the event loop's own thread. The coroutine can only run on that thread, which is blocked waiting for it. Await the coroutine directly when you are already in async code.

Can a threading.Lock deadlock asyncio code?

Yes. If a thread holds the lock while waiting for a coroutine that also needs it, the coroutine blocks the event loop thread on the lock and neither side progresses. Release locks before waiting across the boundary.

How do I detect a sync-to-async call made from the event loop thread?

Check whether asyncio.get_running_loop() returns the loop you are about to wait on; if it does, raise immediately instead of blocking.

Should cross-thread asyncio waits have timeouts?

Always. A deadlock with a timeout becomes an exception with a traceback; without one it is a silent hang. Cancel the waited-on work when the timeout fires.