Inspecting a Live Loop with aiomonitor¶
When an asyncio service is up but stuck — a request that never returns, a background loop that stopped — the question is which task is waiting on what. aiomonitor answers it from inside the running process: started alongside your event loop, it serves a terminal console and a small web UI that list every task, show any task's await stack, and can cancel a task. Measured with aiomonitor 0.7.1 on Python 3.14 against a service with three worker tasks and one task stuck on an asyncio.Event that was never set: the task list showed all five tasks with their names, coroutines and the source line that created each; the stuck task's stack ended at await stuck_event.wait() inside asyncio/locks.py, and its creation stack pointed to the line that started it. Cancelling it through the API returned "Successfully cancelled", and it then appeared in the terminated-task history. Running the monitor without its task-factory hook left task creation within noise of the baseline (12.1 against 14.4 µs, both noisy); with hook_task_factory=True, which records where every task was created, creating and running a task cost 88.9 µs and 50,000 live tasks used about 160 MiB more. This guide sets aiomonitor up safely.
Prerequisites¶
- aiomonitor 0.7+ installed in the service's environment.
- Task inspection basics, from inspecting pending tasks with asyncio.all_tasks.
- The topic overview, Event Loop Debugging & Instrumentation.
1. Start the monitor with the loop¶
aiomonitor runs in a thread of its own and inspects your loop from there, so it keeps working even if your loop is busy. Start it inside main(), around the code that runs the service:
import asyncio
import aiomonitor
async def main():
loop = asyncio.get_running_loop()
with aiomonitor.start_monitor(loop, host="127.0.0.1", port=20101, webui_port=20102):
await run_service()
asyncio.run(main())
The context manager starts a telnet console on port and a web UI on webui_port, and stops both when the block exits. Give tasks names when you create them — asyncio.create_task(coro, name="load-config") — because the task list is only as readable as those names; unnamed tasks appear as Task-1, Task-29 and so on. Connect to the console with python -m aiomonitor.cli (or any telnet client) for an interactive prompt, or open the web UI in a browser.
Verify: the web UI on the chosen port lists your service's tasks, and the console accepts help.
2. List tasks and find the stuck one¶
The task list shows each task's state, name, coroutine, creation location and age. In the console, the command is ps; the web UI's API returns the same data as JSON:
curl -s -X POST http://127.0.0.1:20102/api/live-tasks -d "filter=" | python -m json.tool
Measured on the test service, the list contained main() as the root task, three worker() tasks created at line 12, and load-config running wait_for_config(), created at line 11 and pending for 1 minute 29 seconds. Age is the most useful column when hunting a hang: workers that loop forever are old by design, but a task that should finish in milliseconds and has been pending for minutes is the suspect. The filter parameter narrows the list by name or coroutine, which matters in services with thousands of tasks.
Verify: you can find a task by name and see how long it has been pending.
3. Read the await stack of a task¶
For a suspect task, ask for its stack: in the console where <task_id>, in the web UI the trace page for that task. The stack is the chain of awaits the task is suspended in, innermost last:
File ".../aiomonitor/monitor.py", line 509, in _coro_wrapper
return await coro
File ".../mon_app.py", line 4, in wait_for_config
await stuck_event.wait() # nobody ever sets it
File "/usr/lib/python3.14/asyncio/locks.py", line 213, in wait
await fut
Measured: the stuck task's stack ended in Event.wait, and with the creation hook on, the trace page also showed where the task was created — mon_app.py, line 11, in main. Together they answer "what is it waiting for" and "who started it". A stack ending in a lock, event or queue means the task waits for something in your own code; one ending in a transport or stream read means it waits on the network, and a timeout belongs around it, as covered in Timeouts & Deadlines. The same information is available without aiomonitor from task.get_stack(), as in reading await chains with Task.get_stack; aiomonitor's value is getting it from a process you did not plan to debug.
Verify: for a deliberately hung test task, the stack shown ends at the line it is stuck on.
4. Cancel a task, carefully¶
The console's cancel <task_id>, or a DELETE to the web API, calls task.cancel() on the chosen task:
curl -s -X DELETE "http://127.0.0.1:20102/api/task?task_id=127051645362704"
# {"msg": "Successfully cancelled 127051645362704", "detail": "wait_for_config()"}
Measured: the stuck task disappeared from the live list and appeared in the terminated-task history with its run time. Cancelling is a last resort for unblocking a production process while a fix is prepared: the task's own cleanup runs, but whatever was waiting for its result now receives CancelledError, and cancelling a task that holds a lock or is mid-transaction runs into the issues covered in making database transactions cancellation-safe. Read the stack first, and cancel the innermost task that is actually stuck rather than a parent.
Verify: after cancelling, the task is listed as terminated and the service carries on.
5. Decide on the creation hook, and lock it down¶
With hook_task_factory=True, aiomonitor installs a task factory that wraps every coroutine to record where its task was created and keeps a history of terminated tasks. That is what makes the creation location and the terminated list available — and it is not free:
with aiomonitor.start_monitor(loop, hook_task_factory=False): # production default
...
Measured with 50,000 tasks: creating and running a task took 88.9 µs with the hook against about 12–14 µs without it, and peak memory was about 160 MiB higher while those tasks were alive. In services that create tasks per request, that is a real cost; enable the hook in staging or temporarily while hunting a problem, and rely on task names in production. The console also offers a Python REPL running inside the process — arbitrary code execution by design — so bind only to 127.0.0.1 (the default), never expose the ports beyond the host, and reach them through SSH port forwarding or kubectl port-forward.
Verify: in production the monitor binds to localhost only, and the task-factory hook is off unless deliberately enabled.
Verification¶
aiomonitor is set up well when:
- It starts with the loop in
main(), bound to localhost. - Tasks have meaningful names, so the task list is readable.
- Stacks are read before cancelling, and only the stuck task is cancelled.
- The creation hook is off in production unless deliberately enabled, given its per-task cost.
Diagnostic Hook: when a service stops making progress, read the task list sorted by age before restarting it. A restart destroys the evidence; two minutes in aiomonitor usually shows exactly which task is waiting on which lock, event or socket — and which line created it.
Pitfalls & edge cases¶
- Unnamed tasks. A list of
Task-NNNentries is hard to act on. - The creation hook in production. Measured: 88.9 µs per task and +160 MiB for 50,000 tasks.
- Exposing the console. It includes a Python REPL in the process.
- Cancelling a parent instead of the stuck task. Read the stack first.
Frequently Asked Questions¶
What is aiomonitor?
A library that runs a telnet console and web UI inside an asyncio process, listing tasks, showing their await stacks and creation sites, and cancelling them. Version 0.7.1 was tested here.
How do I find a hung asyncio task in a running service?
With aiomonitor running, list live tasks, look for one pending far longer than it should, and read its stack: the stuck test task's stack ended at await stuck_event.wait().
Does aiomonitor slow down an asyncio application?
Without the task-factory hook, task creation was within noise of the baseline. With hook_task_factory=True, it cost 88.9 µs per task against about 12 to 14 µs.
Is it safe to run aiomonitor in production?
Bound to localhost and with the hook off, yes; its console provides a Python REPL in the process, so never expose its ports.
Related¶
- Event Loop Debugging & Instrumentation — up to the topic overview.
- Tracing awaits with sys.monitoring — timing every await instead of inspecting one.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.