Skip to content

Lazy Async Initialization with a Shared Task

Lazy initialisation — create the connection pool, load the model, fetch the config the first time somebody needs it — is a one-liner in synchronous code and a race in async code. The obvious version checks an attribute, awaits the factory, and stores the result. Every coroutine that arrives during the await sees the attribute still empty and starts its own factory call. In a test with 100 concurrent first callers and a 50 ms factory, the obvious version created 100 separate objects; a version that shares one task among all callers created one. This guide builds that version and handles the three cases that make it non-trivial: failure, cancellation of the first caller, and shutdown.

Prerequisites

1. See why check-then-await races

class Naive:
    def __init__(self) -> None:
        self._pool = None

    async def pool(self):
        if self._pool is None:                 # every early caller sees None...
            self._pool = await create_pool()   # ...and every one of them gets here
        return self._pool

The check and the assignment are separated by an await. In synchronous code with threads you would fix this with a lock; in async code the window is guaranteed to be open, because the await is the point where other tasks run. With 100 concurrent callers, all 100 pass the check before the first factory call returns.

The damage is real: 100 pools means 100× the connections the database expected, and 99 of the pools are overwritten and never closed — leaked sockets, and on many drivers an Unclosed connection warning per leaked connection at garbage collection.

Verify: count factory calls under asyncio.gather(*(obj.pool() for _ in range(100))). Anything above one is the bug.

How 100 callers create 100 pools A sequence of 6 messages between 4 participants. How 100 callers create 100 pools caller 1 lazy holder caller 2 factory _pool is None: call factory create_pool() #1 _pool is still None: call factory create_pool() #2 pool #1 stored pool #2 overwrites it: #1 leaks The await between the check and the store is exactly where every other caller gets in.

2. Share one task among every caller

The fix is to store the in-flight work, not the result. The first caller creates a task; every later caller awaits the same task:

import asyncio
from collections.abc import Awaitable, Callable
from typing import Generic, TypeVar

T = TypeVar("T")


class AsyncLazy(Generic[T]):
    def __init__(self, factory: Callable[[], Awaitable[T]]) -> None:
        self._factory = factory
        self._task: asyncio.Task[T] | None = None

    async def get(self) -> T:
        if self._task is None:
            self._task = asyncio.ensure_future(self._factory())
        return await asyncio.shield(self._task)

The assignment to self._task happens synchronously, before any await, so no other caller can slip between the check and the store. Verified: 100 concurrent callers, 1 factory call, and all 100 received the identical object.

Awaiting a finished task returns its stored result immediately, so after initialisation get() costs one attribute check and one await on a done future — no lock, no re-run.

Verify: the factory call count is 1 and len({id(x) for x in results}) == 1.

3. Shield the shared task from caller cancellation

The asyncio.shield in get() is not decoration. Without it, if the first caller is cancelled — its request times out, the client disconnects — the cancellation propagates into the shared task, and every other caller waiting on it receives a CancelledError they did nothing to cause. The next caller then sees a cancelled task and fails forever.

With the shield, cancelling a caller cancels only that caller's wait; the initialisation continues for everyone else. Verified: cancelling the first caller 10 ms into a 50 ms initialisation left the factory call count at 1, and the next caller received the pool normally.

async def demo(lazy: AsyncLazy) -> None:
    first = asyncio.create_task(lazy.get())
    await asyncio.sleep(0.01)
    first.cancel()                       # only this caller gives up
    pool = await lazy.get()              # initialisation finished anyway

The trade-off: a shielded initialisation cannot be stopped by its callers. If the factory hangs, it hangs for everyone, so put a timeout inside the factory, not around get().

Verify: cancel the first caller mid-initialisation; the factory still runs once and later callers succeed.

The life of an AsyncLazy A flow of 4 stages. The life of an AsyncLazy empty no task yet initialising one shared, shielded task ready done task returns instantly failed reset so the next call retries Storing the task rather than the value is what makes every concurrent caller converge on one initialisation.

4. Retry after failure instead of caching it

As written, a failed initialisation is cached forever: the task holds the exception, and every future get() re-raises it. For a database pool that failed because the database was restarting, that turns a 5-second blip into a permanent outage of the process. Reset on failure:

    async def get(self) -> T:
        if self._task is None or self._failed():
            self._task = asyncio.ensure_future(self._factory())
        return await asyncio.shield(self._task)

    def _failed(self) -> bool:
        t = self._task
        return t is not None and t.done() and (t.cancelled() or t.exception() is not None)

The check runs synchronously before creating a new task, so concurrent callers that arrive after a failure still converge on a single retry. What this does not do is back off: if the factory fails instantly and callers keep arriving, every caller triggers a retry. For dependencies that can be down for a while, put a circuit breaker or a minimum interval between attempts in front of the factory.

Verify: make the factory fail on its first call and succeed on its second; the first batch of callers sees the error and the next caller gets the object.

Lazy initialisation variants compared A grid of 4 rows by 4 columns. Lazy initialisation variants compared approach calls under 100 callers first caller cancelled after a failure check then await 100 n/a retries, races again asyncio.Lock + check 1 lock released, next caller re-runs retries shared task 1 everyone else cancelled failure cached forever shared shielded task + reset 1 init continues next caller retries Only the last row is correct on all three counts; a lock works too but serialises every caller through it.

5. Close what you created

Lazily created resources still need deterministic cleanup, and the cleanup has to handle all three states: never started, still initialising, and ready.

    async def aclose(self, close: Callable[[T], Awaitable[None]]) -> None:
        task, self._task = self._task, None
        if task is None:
            return
        if not task.done():
            task.cancel()
            await asyncio.gather(task, return_exceptions=True)
            return
        if not task.cancelled() and task.exception() is None:
            await close(task.result())

Wire this into the application's lifespan or AsyncExitStack so the pool is closed exactly once at shutdown, as described in managing startup and shutdown with ASGI lifespan. An initialisation still in progress at shutdown is cancelled rather than awaited, so a slow dependency cannot hold the process open.

Verify: call aclose() in each of the three states; no warnings about unclosed resources and no pending tasks remain.

Verification

Lazy initialisation is correct when:

  • Factory calls equal 1 under any number of concurrent first callers.
  • Cancelling a caller never affects other callers or the initialisation.
  • A failure is retried by the next caller rather than cached for the life of the process.
  • Shutdown closes the resource in every state, with no "Task was destroyed but it is pending" message.

Diagnostic Hook: count factory invocations and failures as metrics and log the initialisation duration. More than one invocation per process lifetime without a logged failure means a second code path is creating the resource outside the lazy holder; repeated failures followed by success indicate a dependency that is slow to come up at deploy time.

Pitfalls & edge cases

  • Creating the holder at import time in a module used by several loops. The shared task belongs to the loop it was created on; awaiting it from another loop fails. Create holders per loop or per application instance.
  • functools.cache on an async function. It caches the coroutine object, which can only be awaited once — covered in why functools.lru_cache breaks on coroutines.
  • A timeout around get() instead of inside the factory. With shielding, the caller gives up but the hung initialisation continues and blocks every future caller.
  • Holding the shared task while the factory needs the holder. A factory that calls get() on its own holder awaits itself and deadlocks.

Frequently Asked Questions

How do I lazily initialise a resource once in asyncio?

Store a task for the initialisation rather than its result. The first caller creates the task synchronously; every caller awaits the same task, shielded so one caller's cancellation does not affect the others. Reset the task if it fails so the next caller retries.

Why does my async lazy property create several clients?

Because the check and the assignment are separated by an await. Every caller that arrives while the first factory call is in progress sees the attribute still empty and starts its own call.

Should I use an asyncio.Lock for lazy initialisation?

A lock with a double-check works and produces one call. A shared task is simpler, needs no lock on the hot path, and lets callers be cancelled without affecting the initialisation when combined with asyncio.shield.

Why use asyncio.shield in lazy initialisation?

Without it, cancelling the first caller cancels the shared task, so every other waiting caller fails with CancelledError and the cancelled task would be cached. Shielding confines cancellation to the caller that was cancelled.