Skip to content

Waiting for Readiness with AnyIO TaskGroup.start

Starting a server, a consumer or a connection pool as a background task raises an immediate question: when is it ready? start_soon() returns before the task has done anything, so the next line — connect to the port, publish to the queue — races the startup. AnyIO and trio solve this with TaskGroup.start(): the child receives a task_status object, calls task_status.started(value) when it is ready, and start() returns that value to the parent while the child keeps running in the group. In a test, a server bound to port 0 handed its real port back through started(), and the parent connected immediately and received b'hello'. Two edge cases behaved exactly as you would want: a child that raised before calling started() raised in the caller of start() (OSError: bind failed) without killing the group, and a child that returned without calling it produced RuntimeError: Child exited without calling task_status.started().

Prerequisites

1. Write a startable task

A startable task accepts a keyword-only task_status defaulting to anyio.TASK_STATUS_IGNORED, so it can also be run with start_soon or awaited directly:

import anyio
from anyio.abc import SocketAttribute, TaskStatus


async def serve(*, task_status: TaskStatus[int] = anyio.TASK_STATUS_IGNORED) -> None:
    listener = await anyio.create_tcp_listener(local_host="127.0.0.1", local_port=0)
    port = listener.extra(SocketAttribute.local_port)
    task_status.started(port)                     # ready: the socket is listening

    async def handle(stream):
        async with stream:
            await stream.send(b"hello")

    await listener.serve(handle)                  # keeps running inside the task group

The value passed to started() is whatever the parent needs to use the task: here the OS-assigned port, elsewhere a client object, a URL, or nothing at all. Call started() at the precise point after which the task is usable — after bind and listen, after a consumer subscribed, after a pool connected — not at the top of the function.

Verify: await tg.start(serve) returns an int, and a connection to that port succeeds immediately.

start() blocks until the child says it is ready A sequence of 6 messages between 3 participants. start() blocks until the child says it is ready parent task group child serve() await tg.start(serve) run child bind, listen on port 0 task_status.started(port) start() returns port keeps serving in the group The parent cannot race the startup, because start() only returns once the child is ready.

2. Use start() in the parent

async def main() -> None:
    async with anyio.create_task_group() as tg:
        port = await tg.start(serve)
        async with await anyio.connect_tcp("127.0.0.1", port) as conn:
            print(await conn.receive())               # b'hello'
        tg.cancel_scope.cancel()                       # shut the server down


anyio.run(main)

Startup becomes sequential where it needs to be and concurrent everywhere else: start the database pool, then the consumer that needs it, then the HTTP server that needs both, each await tg.start(...) returning only when that layer is up. Compared with sleeping "long enough", it is both faster and correct; compared with an Event per component, it carries a value and has error handling built in.

Verify: start three dependent components in order; the logs show each one ready before the next begins.

3. Let startup failures reach the caller

The part start() gets right that ad-hoc readiness events usually get wrong is failure. If the child raises before calling started(), the exception is raised from await tg.start(...) in the parent — not into the task group:

async def flaky_start(*, task_status=anyio.TASK_STATUS_IGNORED) -> None:
    raise OSError("bind failed")


async def main() -> None:
    async with anyio.create_task_group() as tg:
        try:
            await tg.start(flaky_start)
        except OSError as exc:
            log.error("component failed to start: %s", exc)
            # the group is still alive; decide whether to retry, degrade or exit

Verified: the OSError was caught around start() and the group kept running. Errors after started() follow normal task-group rules — they cancel the group and propagate out of the async with. That split matches what callers need: a startup failure is the caller's decision to handle; a runtime failure of an established component is a crash of the whole unit.

A child that returns without ever calling started() is a programming error, and start() raises RuntimeError: Child exited without calling task_status.started() rather than hanging.

Verify: inject a startup failure; the caller handles it, and other tasks in the group are unaffected.

Where a child's exception goes A grid of 3 rows by 3 columns. Where a child's exception goes child does before started() after started() raises an exception raised from await start() cancels the task group returns normally RuntimeError from start() task simply ends is cancelled CancelledError in the caller group carries on Startup failures belong to the caller; runtime failures belong to the group.

4. Bound how long startup may take

start() waits as long as the child takes to become ready. Put a deadline on it, so a dependency that never comes up fails startup instead of hanging the process:

async def main() -> None:
    async with anyio.create_task_group() as tg:
        with anyio.fail_after(10):
            pool = await tg.start(run_db_pool)
        with anyio.fail_after(5):
            await tg.start(run_consumer, pool)
        await tg.start(run_http, pool)

fail_after raises TimeoutError when the deadline passes and cancels the child that was starting. The other components already started are untouched until the exception leaves the task group. The deadline belongs around start(), not inside the child, because only the caller knows how long it is willing to wait. The scope semantics are covered in using AnyIO cancel scopes and move_on_after.

Verify: make the pool hang during connect; startup fails after 10 s with a TimeoutError naming the pool step in your logs.

5. Replicate it on plain asyncio

asyncio.TaskGroup has no start(). The same contract takes a future and a wrapper:

import asyncio


async def start(tg: asyncio.TaskGroup, fn, *args):
    loop = asyncio.get_running_loop()
    ready: asyncio.Future = loop.create_future()

    async def runner():
        try:
            await fn(*args, started=lambda v=None: ready.done() or ready.set_result(v))
        except BaseException as exc:
            if not ready.done():
                ready.set_exception(exc)      # startup failure goes to the caller...
                return                        # ...not to the group
            raise
        if not ready.done():
            ready.set_exception(RuntimeError("child exited without calling started()"))

    tg.create_task(runner())
    return await ready

The child calls started(value) instead of task_status.started(value). Catching BaseException routes a startup-time cancellation to the caller too; for production use, consider whether that is what you want, or switch to AnyIO and get the semantics without maintaining them. The broader comparison is in comparing trio nurseries and asyncio TaskGroup.

Verify: the three behaviours from section 3 — value, startup error, missing started() — all hold for the asyncio helper.

Ordered startup with start() A flow of 4 stages. Ordered startup with start() start(run_db_pool) returns the pool start(run_consumer, pool) subscribed start(run_http, pool) listening mark ready every layer up Each await returns only when that layer is usable, so readiness is exact and failures name their step.

Verification

Readiness handling is correct when:

  • Every background component is started with start() and calls started() exactly when usable.
  • Startup failures are handled by the caller, and runtime failures stop the group.
  • Each start() has a deadline.
  • Readiness probes flip only after the last start() returns.

Diagnostic Hook: log each component's time from start() to started() and export it as a startup-duration metric per component. Over deploys this shows which dependency dominates cold start; a component whose startup time creeps up is usually waiting on a slower network path or a larger warm-up than before.

Pitfalls & edge cases

  • Calling started() too early. Before listen() or before a subscription is confirmed, the parent races the child again.
  • Positional task_status. It must be keyword-only, or start_soon passes arguments into it.
  • Calling started() twice. The second call raises; guard it if several paths can signal readiness.
  • Long work before started() without a deadline. The caller hangs as long as the child does.

Frequently Asked Questions

What does task_status.started() do in AnyIO?

It signals to the parent that called TaskGroup.start() that the child is ready, optionally passing a value. start() then returns that value while the child keeps running inside the task group.

What happens if a task fails before calling started()?

The exception is raised from await tg.start() in the caller, and the task group is not cancelled. If the child returns without calling started(), start() raises RuntimeError.

Does asyncio.TaskGroup have a start method?

No. You can replicate it with a future that the child resolves when ready and that receives any exception raised before readiness, or use AnyIO's task groups on top of asyncio.

How do I wait for a server task to be listening before connecting?

Have the server call task_status.started(port) after binding and listening, and start it with port = await tg.start(server). The connection can then be made immediately.