Skip to content

Logging Task Lifecycles with a Custom Task Factory

Every task in an asyncio program — yours, your framework's, every library's — is created through one hook: the loop's task factory. Installing a custom factory with loop.set_task_factory() gives you a single place to count tasks, time their lifetimes, record where they were created, and report failures, without touching any call site. Done carelessly it slows every task and breaks libraries that subclass Task; done properly it is cheap. Measured on Python 3.14, an instrumented factory that counted, timed and attached a completion callback raised the cost of creating and running a trivial task from 1.96 µs to 2.82 µs — under a microsecond per task — and accounted for all 100,000 tasks in the benchmark. This guide builds that factory, makes it compatible with the eager task factory, and turns its data into metrics.

Prerequisites

1. Install a factory that wraps Task creation

A task factory is a callable factory(loop, coro, **kwargs) returning a task. Keep the default construction and decorate the result:

import asyncio
import time
from collections import Counter

task_stats: Counter[str] = Counter()


def _on_done(task: asyncio.Task) -> None:
    task_stats["finished"] += 1
    if task.cancelled():
        task_stats["cancelled"] += 1
    elif task.exception() is not None:
        task_stats["failed"] += 1
    lifetime_hist.observe(time.perf_counter() - task._created_at)


def instrumented_factory(loop, coro, **kwargs) -> asyncio.Task:
    task = asyncio.Task(coro, loop=loop, **kwargs)   # kwargs carries name= and context=
    task._created_at = time.perf_counter()
    task_stats["created"] += 1
    task.add_done_callback(_on_done)
    return task


async def main() -> None:
    asyncio.get_running_loop().set_task_factory(instrumented_factory)
    await serve()

Pass **kwargs through untouched: create_task forwards name= and context= (and eager_start= on 3.14) to the factory, and dropping them silently loses task names and breaks context propagation. Install the factory as the first thing main() does, so tasks created during startup are counted too.

Note that calling task.exception() in the callback marks failures as retrieved; that suppresses the "Task exception was never retrieved" log for every task in the process. If you rely on that log, log failures here explicitly, with the task name, instead.

Verify: after a workload, created - finished equals the number of tasks still pending, as reported by len(asyncio.all_tasks()).

One hook sees every task in the process A flow of 4 stages. One hook sees every task in the process create_task() yours or a library's loop task factory builds the Task stamp + count created_at, created += 1 done-callback outcome and lifetime Because every task passes through the factory, instrumentation needs no changes at call sites.

2. Keep the overhead small

Everything the factory does runs for every task, so measure it. The benchmark: create and gather 100,000 trivial tasks in batches of 1,000, with and without the factory:

async def noop() -> None:
    pass


async def bench(n: int = 100_000) -> float:
    t = time.perf_counter()
    for _ in range(n // 1000):
        await asyncio.gather(*(asyncio.create_task(noop()) for _ in range(1000)))
    return (time.perf_counter() - t) / n * 1e6      # µs per task

Results: 1.96 µs per task without the factory, 2.82 µs with counting, timing and a done-callback. Real tasks do far more than nothing, so in practice the overhead disappears into noise. What does not disappear is anything expensive per task: capturing a full stack with traceback.extract_stack() costs tens of microseconds, and logging a line per task at INFO turns a busy service's log volume into its main workload. Capture a single caller frame, as in reading await chains with task.get_stack, and aggregate into metrics rather than logging individual tasks.

Verify: run the benchmark with your factory; the added cost per task should be around a microsecond.

Cost per trivial task, with and without instrumentation 2 horizontal bars comparing instrumented factory with the others. Cost per trivial task, with and without instrumentation instrumented factory 2.82 µs per task default factory 1.96 µs per task 100,000 no-op tasks gathered in batches of 1,000; Python 3.14 on Linux. Counting and timing every task costs under a microsecond; full stack capture would cost far more.

3. Stay compatible with eager tasks and Task subclasses

There is only one task factory slot per loop. If you also want the eager task factory from Python 3.12, compose them: asyncio.create_eager_task_factory(custom_task_constructor) builds an eager factory around any Task-compatible constructor, and a Task subclass is the cleanest place for the instrumentation:

class InstrumentedTask(asyncio.Task):
    def __init__(self, coro, *, loop=None, **kwargs) -> None:
        self._created_at = time.perf_counter()
        task_stats["created"] += 1
        super().__init__(coro, loop=loop, **kwargs)
        self.add_done_callback(_on_done)


loop = asyncio.get_running_loop()
loop.set_task_factory(asyncio.create_eager_task_factory(InstrumentedTask))

Verified: tasks created afterwards were InstrumentedTask instances, and a task whose coroutine finished without awaiting was already done when create_task returned — eager execution preserved. Set the counters before super().__init__, because with eager start the coroutine may run, finish and fire callbacks inside the constructor. How eager tasks change ordering is covered in using the eager task factory in Python 3.12.

Check before installing that nothing else already set a factory — some frameworks and tracing libraries do. loop.get_task_factory() returns the current one; wrap it rather than replacing it.

Verify: type(asyncio.create_task(x())) is your class, and loop.get_task_factory() was None (or was wrapped) before you installed yours.

4. Turn the counts into metrics

Raw counters become useful when exported with a little structure. Group by a name prefix, so "which kind of task is leaking" is answerable:

from collections import defaultdict

by_kind: dict[str, Counter] = defaultdict(Counter)


def kind_of(task: asyncio.Task) -> str:
    name = task.get_name()
    return name.split(":", 1)[0] if ":" in name else "unnamed"


def _on_done(task: asyncio.Task) -> None:
    k = kind_of(task)
    by_kind[k]["finished"] += 1
    outcome = "cancelled" if task.cancelled() else "failed" if task.exception() else "ok"
    by_kind[k][outcome] += 1


def export(metrics) -> None:
    for k, c in by_kind.items():
        metrics.gauge("tasks_alive", c["created"] - c["finished"], kind=k)
        metrics.counter("tasks_failed_total", c["failed"], kind=k)

Naming tasks request:…, refresh:…, consumer:… at creation keeps the label set small and meaningful; unnamed tasks from libraries fall into one bucket. Exporting these to Prometheus follows the patterns in exporting Prometheus metrics from asyncio.

Verify: under steady load, tasks_alive per kind is flat; it rises only with concurrency, not with time.

What the factory's numbers tell you A grid of 4 rows by 3 columns. What the factory's numbers tell you signal usual cause look at alive count grows with uptime tasks never finish that kind's await chains failed rate rises a dependency is erroring failure logs by kind cancelled rate rises clients or timeouts giving up latency upstream lifetime p99 jumps something they await is slow pools, locks, downstreams The factory's counters point at a kind of task; await-chain dumps then point at a line.

5. Remove or gate it cleanly

Instrumentation should be switchable at runtime, because it runs on the hottest path in the process. Read a flag once at startup, and restore the previous factory when shutting down tests:

import contextlib
import os


@contextlib.contextmanager
def task_instrumentation(loop: asyncio.AbstractEventLoop):
    if os.environ.get("TASK_METRICS", "1") != "1":
        yield
        return
    previous = loop.get_task_factory()
    loop.set_task_factory(instrumented_factory)
    try:
        yield
    finally:
        loop.set_task_factory(previous)

In tests, the same factory doubles as a leak detector: assert that created == finished at the end of each test, the approach expanded in detecting leaked tasks in tests.

Verify: with TASK_METRICS=0, type(asyncio.create_task(x())) is plain asyncio.Task and benchmarks match the uninstrumented baseline.

Verification

The factory is production-ready when:

  • Every task is counted — created minus finished equals len(asyncio.all_tasks()).
  • Overhead is around a microsecond per task, measured in your environment.
  • Names and contexts are preserved, and eager execution still works if you use it.
  • It can be switched off without a code change.

Diagnostic Hook: alert on tasks_alive per kind growing for more than an hour at constant traffic, and on failed-to-created ratio per kind above a threshold. The first catches leaks long before memory alerts fire; the second catches background work failing silently — the class of bug from handling exceptions in fire-and-forget tasks.

Pitfalls & edge cases

  • Dropping **kwargs. Task names, contexts and eager flags silently disappear.
  • Replacing another library's factory. Check get_task_factory() first and wrap it.
  • Expensive work per task. Full stack captures and per-task log lines dominate CPU in task-heavy services.
  • Counting before super().__init__ raises. If construction fails, decrement or count after success, or created drifts upward.

Frequently Asked Questions

What is an asyncio task factory?

A callable installed with loop.set_task_factory that the loop uses to create every Task, receiving the loop, the coroutine and keyword arguments such as name and context. It is the single place where all task creation in a process can be observed.

How much overhead does a custom task factory add?

One that counts tasks, records a timestamp and attaches a done-callback raised the cost of a trivial task from 1.96 µs to 2.82 µs on Python 3.14. Real tasks do much more work, so the relative cost is small.

Can I combine a custom task factory with eager tasks?

Yes. Subclass asyncio.Task with your instrumentation and pass it to asyncio.create_eager_task_factory, then install the result with set_task_factory.

How do I count asyncio tasks by type?

Name tasks with a prefix at creation, such as request: or refresh:, and in the factory's done-callback increment counters keyed by that prefix. Export created minus finished per prefix as a gauge.