Skip to content

Dumping Stacks of a Hung asyncio Program

"The service is hung" means one of two very different things in asyncio, and the first job is to tell them apart. Either the loop is blocked — one callback or task step is stuck in a synchronous call, so nothing else runs — or the loop is idle and every task is waiting on something that will never happen: a lock nobody releases, a future nobody resolves, a socket that never answers. A thread dump distinguishes them in one look: a blocked loop shows your code on the main thread's stack, an idle one shows selectors.select. In the idle case the thread dump is useless — it shows the selector and nothing else — and you need a task dump. On Python 3.14, asyncio.print_call_graph() produced the full await chain of every task, including which task was awaiting which, from a signal handler in a running service. This guide sets up both dumps before you need them.

Prerequisites

1. Register faulthandler for thread stacks

faulthandler dumps the Python stack of every thread from a signal handler implemented in C, so it works even when the event loop thread is completely stuck in a blocking call. Register it at startup:

import faulthandler
import signal

faulthandler.register(signal.SIGUSR1, all_threads=True)    # kill -USR1 <pid>

On an idle loop, the dump of the main thread looks like this — the loop sitting in select, waiting for I/O or a timer:

Current thread 0x000073a8c183f200 [python3] (most recent call first):
  File "/usr/lib/python3.14/selectors.py", line 452 in select
  File "/usr/lib/python3.14/asyncio/base_events.py", line 2019 in _run_once
  File "/usr/lib/python3.14/asyncio/base_events.py", line 677 in run_forever
  ...

If instead the top frames are your code — requests.get, time.sleep, a tight loop in a parser — the loop is blocked, and that frame is the culprit. faulthandler.dump_traceback_later(timeout, repeat=True) does the same on a timer, which is useful for capturing a hang that happens at 3 a.m.

Verify: kill -USR1 a running service; its stderr shows one stack per thread without the process exiting.

Blocked loop or idle loop? A decision on What is on top of the loop thread's stack with 3 outcomes. Blocked loop or idle loop? What is on top of the loop thread's stack? your code or a sync library blocked loop that frame is the bug selectors.select idle loop dump the tasks next a lock or queue in a thread thread deadlock check executor threads One thread dump splits the problem in two; the idle case needs a task dump to go further.

2. Add an on-demand task dump

When the loop is idle, the information is in the tasks: which ones exist, and what each is awaiting. Python 3.14's asyncio.print_call_graph() prints a task's full chain of awaits and the tasks awaiting it. Hook it to a signal on the loop:

import asyncio
import signal
import sys


def dump_tasks() -> None:
    tasks = asyncio.all_tasks()
    print(f"=== {len(tasks)} tasks ===", file=sys.stderr)
    for task in tasks:
        asyncio.print_call_graph(task, file=sys.stderr, limit=8)


async def main() -> None:
    asyncio.get_running_loop().add_signal_handler(signal.SIGUSR2, dump_tasks)
    await serve()

Real output from a service with two request tasks stuck in a slow query under a TaskGroup:

=== 3 tasks ===
* Task(name='request-1', id=0x73a8c001c050)
  + Call stack:
  |   File '/usr/lib/python3.14/asyncio/tasks.py', line 702, in async sleep()
  |   File 'service.py', line 7, in async db_query()
  |   File 'service.py', line 8, in async handle()
  + Awaited by:
    * Task(name='Task-1', id=0x73a8c00205f0)
      + Call stack:
      |   File '/usr/lib/python3.14/asyncio/taskgroups.py', line 72, in async TaskGroup.__aexit__()
      |   File 'service.py', line 11, in async main()

The "Awaited by" section is what makes this more than a stack: it shows the structure — which parent is blocked waiting on which child. A signal handler added with add_signal_handler runs on the loop, so it needs the loop to be idle, not blocked — which is exactly the case this dump is for.

Verify: kill -USR2 the service; every task appears with its innermost await on top.

3. Fall back to get_stack on older Pythons

Before 3.14, task.get_stack() is the built-in option, and it has a big limitation: for a task running a coroutine it returns only the task's own coroutine frame, not the chain below it. Verified on 3.14 for comparison: for a task suspended in top → middle → leaf → sleep, get_stack() returned just ['top'], while the call graph returned all four. Walk the await chain yourself to recover the rest:

def await_chain(task: asyncio.Task) -> list[str]:
    frames = []
    coro = task.get_coro()
    while coro is not None:
        frame = getattr(coro, "cr_frame", None) or getattr(coro, "ag_frame", None)
        if frame is not None:
            frames.append(f"{frame.f_code.co_filename}:{frame.f_lineno} {frame.f_code.co_name}")
        coro = getattr(coro, "cr_await", None) or getattr(coro, "ag_await", None)
    return frames


def dump_tasks_legacy() -> None:
    for t in asyncio.all_tasks():
        print(t.get_name(), "\n  " + "\n  ".join(await_chain(t)), file=sys.stderr)

Following cr_await walks from each coroutine to the one it is awaiting until it reaches a future, which is where the task is actually parked. It does not show the "awaited by" relationships, but naming tasks makes those easy to infer, as described in naming and tracking tasks for observability. The details of what get_stack does and does not show are in reading await chains with task.get_stack.

Verify: the legacy dump shows the innermost coroutine (the one calling sleep, read, acquire) for each task.

Which dump answers which question A grid of 4 rows by 4 columns. Which dump answers which question tool shows loop blocked? setup faulthandler + SIGUSR1 every thread's stack works one line at startup print_call_graph via signal every task's await chain needs an idle loop 3.14+ cr_await walk via signal innermost await per task needs an idle loop any version py-spy dump thread stacks, from outside works install, ptrace rights Thread dumps diagnose a blocked loop; task dumps diagnose an idle one.

4. Dump from outside the process

When you could not prepare the process — a production incident in code without signal handlers — external tools read the stacks via ptrace. py-spy dump --pid <pid> prints every thread's Python stack, like faulthandler, without touching the process; covered in profiling asyncio applications with py-spy.

Python 3.14 adds python -m asyncio ps <pid> and python -m asyncio pstree <pid>, which read tasks from outside, giving the call-graph view without in-process setup. They need permission to read the target's memory — the same ptrace rights as a debugger. Under Linux's default Yama setting (/proc/sys/kernel/yama/ptrace_scope = 1), only a parent process may do that, and an unprivileged ps run from another shell fails with the misleading Failed to find the PyRuntime section in process … on Linux platform. Verified on 3.14.4: the same command printed the full task table once the target allowed tracing. Plan the permission (a debug sidecar with SYS_PTRACE, or a target that opts in) before you need it, and keep the in-process task dump for environments where you cannot.

py-spy dump --pid 12345                      # thread stacks, works on most builds
python3.14 -m asyncio pstree 12345           # task tree, needs ptrace rights to the target

In containers, both need SYS_PTRACE or a debug sidecar sharing the process namespace; plan that before the incident.

Verify: run each tool against a staging process now, so you know which ones work on your build and image.

5. Make hangs leave evidence on their own

The best hang dump is one taken automatically before anyone notices. Combine a watchdog with the dumps above:

import faulthandler
import time
import threading


def start_watchdog(loop: asyncio.AbstractEventLoop, stall_s: float = 5.0) -> None:
    last = [time.monotonic()]

    async def heartbeat() -> None:
        while True:
            last[0] = time.monotonic()
            await asyncio.sleep(1)

    def watch() -> None:
        while True:
            time.sleep(stall_s)
            if time.monotonic() - last[0] > stall_s:
                faulthandler.dump_traceback(all_threads=True)       # loop is blocked: show why

    loop.create_task(heartbeat(), name="watchdog-heartbeat")
    threading.Thread(target=watch, name="watchdog", daemon=True).start()

A heartbeat that stops updating means the loop is blocked; the watchdog thread, which is not on the loop, dumps the thread stacks at that moment and catches the blocking frame in the act. For idle-loop hangs, schedule the task dump when request latency or queue age crosses a threshold. The lag measurement underneath this is described in measuring event loop lag in production.

Verify: add a time.sleep(10) to a handler; within five seconds the logs contain a thread dump with that line on top.

A watchdog that captures the blocking frame A flow of 4 stages. A watchdog that captures the blocking frame heartbeat task updates timestamp each second watchdog thread off the loop, checks it stale timestamp the loop is blocked faulthandler dump the blocking frame, live Because the watchdog is not on the loop, it can observe the loop being stuck.

Verification

You are ready for the next hang when:

  • SIGUSR1 produces thread stacks in every deployed process.
  • SIGUSR2 (or an admin endpoint) produces a task dump with await chains.
  • A watchdog dumps automatically when the loop stops making progress.
  • External tools — py-spy, and on 3.14 asyncio ps/pstree — have been tried against a staging process with the ptrace rights they need.

Diagnostic Hook: ship dumps to the same place as your logs with the process id and a timestamp, and count them as a metric. A watchdog dump is always worth an alert; a rising count of tasks in task dumps that share the same innermost await (the same lock, the same pool acquire) points straight at the contended resource.

Pitfalls & edge cases

  • Expecting a task dump from a blocked loop. Signal handlers added with add_signal_handler run on the loop; use faulthandler for that case.
  • Huge dumps. Services with tens of thousands of tasks produce enormous output; summarise by innermost frame first, then print a few examples of each.
  • Windows. add_signal_handler is unavailable and SIGUSR1 does not exist; expose the dump through an admin endpoint or a file trigger instead.
  • Signals swallowed by the process manager. Some supervisors forward only SIGTERM; test that SIGUSR1 reaches the process in your deployment.

Frequently Asked Questions

How do I see what a hung asyncio program is doing?

First dump thread stacks with faulthandler or py-spy. If the main thread is in your code, the loop is blocked by that call. If it is in selectors.select, the loop is idle and you need a task dump — asyncio.print_call_graph on 3.14, or a cr_await walk on earlier versions.

Why does task.get_stack only show one frame?

For a task running a coroutine, get_stack returns the task's own coroutine frame rather than the whole chain of awaits below it. Follow cr_await from coroutine to coroutine, or use asyncio.capture_call_graph on Python 3.14.

What do python -m asyncio ps and pstree do?

Added in Python 3.14, they list the asyncio tasks of another running process and their await relationships, by reading its memory like a debugger. They need ptrace permission to the target; without it they fail with a misleading 'Failed to find the PyRuntime section' error, so keep an in-process dump too.

Can faulthandler dump a process whose event loop is blocked?

Yes. faulthandler's handler is implemented in C and runs even while the main thread is stuck in a blocking call, which is exactly when an on-loop handler cannot run.