Logging Task Lifecycles with a Custom Task Factory¶
Every task in an asyncio program — yours, your framework's, every library's — is created through one hook: the loop's task factory. Installing a custom factory with loop.set_task_factory() gives you a single place to count tasks, time their lifetimes, record where they were created, and report failures, without touching any call site. Done carelessly it slows every task and breaks libraries that subclass Task; done properly it is cheap. Measured on Python 3.14, an instrumented factory that counted, timed and attached a completion callback raised the cost of creating and running a trivial task from 1.96 µs to 2.82 µs — under a microsecond per task — and accounted for all 100,000 tasks in the benchmark. This guide builds that factory, makes it compatible with the eager task factory, and turns its data into metrics.
Prerequisites¶
- Python 3.12+ for
asyncio.create_eager_task_factory; the plain factory works on 3.11. - Task naming, from naming and tracking tasks for observability.
- Done-callback rules, from using add_done_callback without losing exceptions.
1. Install a factory that wraps Task creation¶
A task factory is a callable factory(loop, coro, **kwargs) returning a task. Keep the default construction and decorate the result:
import asyncio
import time
from collections import Counter
task_stats: Counter[str] = Counter()
def _on_done(task: asyncio.Task) -> None:
task_stats["finished"] += 1
if task.cancelled():
task_stats["cancelled"] += 1
elif task.exception() is not None:
task_stats["failed"] += 1
lifetime_hist.observe(time.perf_counter() - task._created_at)
def instrumented_factory(loop, coro, **kwargs) -> asyncio.Task:
task = asyncio.Task(coro, loop=loop, **kwargs) # kwargs carries name= and context=
task._created_at = time.perf_counter()
task_stats["created"] += 1
task.add_done_callback(_on_done)
return task
async def main() -> None:
asyncio.get_running_loop().set_task_factory(instrumented_factory)
await serve()
Pass **kwargs through untouched: create_task forwards name= and context= (and eager_start= on 3.14) to the factory, and dropping them silently loses task names and breaks context propagation. Install the factory as the first thing main() does, so tasks created during startup are counted too.
Note that calling task.exception() in the callback marks failures as retrieved; that suppresses the "Task exception was never retrieved" log for every task in the process. If you rely on that log, log failures here explicitly, with the task name, instead.
Verify: after a workload, created - finished equals the number of tasks still pending, as reported by len(asyncio.all_tasks()).
2. Keep the overhead small¶
Everything the factory does runs for every task, so measure it. The benchmark: create and gather 100,000 trivial tasks in batches of 1,000, with and without the factory:
async def noop() -> None:
pass
async def bench(n: int = 100_000) -> float:
t = time.perf_counter()
for _ in range(n // 1000):
await asyncio.gather(*(asyncio.create_task(noop()) for _ in range(1000)))
return (time.perf_counter() - t) / n * 1e6 # µs per task
Results: 1.96 µs per task without the factory, 2.82 µs with counting, timing and a done-callback. Real tasks do far more than nothing, so in practice the overhead disappears into noise. What does not disappear is anything expensive per task: capturing a full stack with traceback.extract_stack() costs tens of microseconds, and logging a line per task at INFO turns a busy service's log volume into its main workload. Capture a single caller frame, as in reading await chains with task.get_stack, and aggregate into metrics rather than logging individual tasks.
Verify: run the benchmark with your factory; the added cost per task should be around a microsecond.
3. Stay compatible with eager tasks and Task subclasses¶
There is only one task factory slot per loop. If you also want the eager task factory from Python 3.12, compose them: asyncio.create_eager_task_factory(custom_task_constructor) builds an eager factory around any Task-compatible constructor, and a Task subclass is the cleanest place for the instrumentation:
class InstrumentedTask(asyncio.Task):
def __init__(self, coro, *, loop=None, **kwargs) -> None:
self._created_at = time.perf_counter()
task_stats["created"] += 1
super().__init__(coro, loop=loop, **kwargs)
self.add_done_callback(_on_done)
loop = asyncio.get_running_loop()
loop.set_task_factory(asyncio.create_eager_task_factory(InstrumentedTask))
Verified: tasks created afterwards were InstrumentedTask instances, and a task whose coroutine finished without awaiting was already done when create_task returned — eager execution preserved. Set the counters before super().__init__, because with eager start the coroutine may run, finish and fire callbacks inside the constructor. How eager tasks change ordering is covered in using the eager task factory in Python 3.12.
Check before installing that nothing else already set a factory — some frameworks and tracing libraries do. loop.get_task_factory() returns the current one; wrap it rather than replacing it.
Verify: type(asyncio.create_task(x())) is your class, and loop.get_task_factory() was None (or was wrapped) before you installed yours.
4. Turn the counts into metrics¶
Raw counters become useful when exported with a little structure. Group by a name prefix, so "which kind of task is leaking" is answerable:
from collections import defaultdict
by_kind: dict[str, Counter] = defaultdict(Counter)
def kind_of(task: asyncio.Task) -> str:
name = task.get_name()
return name.split(":", 1)[0] if ":" in name else "unnamed"
def _on_done(task: asyncio.Task) -> None:
k = kind_of(task)
by_kind[k]["finished"] += 1
outcome = "cancelled" if task.cancelled() else "failed" if task.exception() else "ok"
by_kind[k][outcome] += 1
def export(metrics) -> None:
for k, c in by_kind.items():
metrics.gauge("tasks_alive", c["created"] - c["finished"], kind=k)
metrics.counter("tasks_failed_total", c["failed"], kind=k)
Naming tasks request:…, refresh:…, consumer:… at creation keeps the label set small and meaningful; unnamed tasks from libraries fall into one bucket. Exporting these to Prometheus follows the patterns in exporting Prometheus metrics from asyncio.
Verify: under steady load, tasks_alive per kind is flat; it rises only with concurrency, not with time.
5. Remove or gate it cleanly¶
Instrumentation should be switchable at runtime, because it runs on the hottest path in the process. Read a flag once at startup, and restore the previous factory when shutting down tests:
import contextlib
import os
@contextlib.contextmanager
def task_instrumentation(loop: asyncio.AbstractEventLoop):
if os.environ.get("TASK_METRICS", "1") != "1":
yield
return
previous = loop.get_task_factory()
loop.set_task_factory(instrumented_factory)
try:
yield
finally:
loop.set_task_factory(previous)
In tests, the same factory doubles as a leak detector: assert that created == finished at the end of each test, the approach expanded in detecting leaked tasks in tests.
Verify: with TASK_METRICS=0, type(asyncio.create_task(x())) is plain asyncio.Task and benchmarks match the uninstrumented baseline.
Verification¶
The factory is production-ready when:
- Every task is counted — created minus finished equals
len(asyncio.all_tasks()). - Overhead is around a microsecond per task, measured in your environment.
- Names and contexts are preserved, and eager execution still works if you use it.
- It can be switched off without a code change.
Diagnostic Hook: alert on tasks_alive per kind growing for more than an hour at constant traffic, and on failed-to-created ratio per kind above a threshold. The first catches leaks long before memory alerts fire; the second catches background work failing silently — the class of bug from handling exceptions in fire-and-forget tasks.
Pitfalls & edge cases¶
- Dropping
**kwargs. Task names, contexts and eager flags silently disappear. - Replacing another library's factory. Check
get_task_factory()first and wrap it. - Expensive work per task. Full stack captures and per-task log lines dominate CPU in task-heavy services.
- Counting before
super().__init__raises. If construction fails, decrement or count after success, orcreateddrifts upward.
Frequently Asked Questions¶
What is an asyncio task factory?
A callable installed with loop.set_task_factory that the loop uses to create every Task, receiving the loop, the coroutine and keyword arguments such as name and context. It is the single place where all task creation in a process can be observed.
How much overhead does a custom task factory add?
One that counts tasks, records a timestamp and attaches a done-callback raised the cost of a trivial task from 1.96 µs to 2.82 µs on Python 3.14. Real tasks do much more work, so the relative cost is small.
Can I combine a custom task factory with eager tasks?
Yes. Subclass asyncio.Task with your instrumentation and pass it to asyncio.create_eager_task_factory, then install the result with set_task_factory.
How do I count asyncio tasks by type?
Name tasks with a prefix at creation, such as request: or refresh:, and in the factory's done-callback increment counters keyed by that prefix. Export created minus finished per prefix as a gauge.
Related¶
- Event Loop Debugging & Instrumentation — up to the topic overview.
- Finding never-retrieved task exceptions — the failure-reporting side of the same hook.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.