Lazy Async Initialization with a Shared Task¶
Lazy initialisation — create the connection pool, load the model, fetch the config the first time somebody needs it — is a one-liner in synchronous code and a race in async code. The obvious version checks an attribute, awaits the factory, and stores the result. Every coroutine that arrives during the await sees the attribute still empty and starts its own factory call. In a test with 100 concurrent first callers and a 50 ms factory, the obvious version created 100 separate objects; a version that shares one task among all callers created one. This guide builds that version and handles the three cases that make it non-trivial: failure, cancellation of the first caller, and shutdown.
Prerequisites¶
- Python 3.11+, stdlib only.
- Why
awaitopens a window for other tasks, from Coroutine Design Patterns. - Single-flight, the per-key generalisation of this pattern, from implementing the single-flight pattern.
1. See why check-then-await races¶
class Naive:
def __init__(self) -> None:
self._pool = None
async def pool(self):
if self._pool is None: # every early caller sees None...
self._pool = await create_pool() # ...and every one of them gets here
return self._pool
The check and the assignment are separated by an await. In synchronous code with threads you would fix this with a lock; in async code the window is guaranteed to be open, because the await is the point where other tasks run. With 100 concurrent callers, all 100 pass the check before the first factory call returns.
The damage is real: 100 pools means 100× the connections the database expected, and 99 of the pools are overwritten and never closed — leaked sockets, and on many drivers an Unclosed connection warning per leaked connection at garbage collection.
Verify: count factory calls under asyncio.gather(*(obj.pool() for _ in range(100))). Anything above one is the bug.
2. Share one task among every caller¶
The fix is to store the in-flight work, not the result. The first caller creates a task; every later caller awaits the same task:
import asyncio
from collections.abc import Awaitable, Callable
from typing import Generic, TypeVar
T = TypeVar("T")
class AsyncLazy(Generic[T]):
def __init__(self, factory: Callable[[], Awaitable[T]]) -> None:
self._factory = factory
self._task: asyncio.Task[T] | None = None
async def get(self) -> T:
if self._task is None:
self._task = asyncio.ensure_future(self._factory())
return await asyncio.shield(self._task)
The assignment to self._task happens synchronously, before any await, so no other caller can slip between the check and the store. Verified: 100 concurrent callers, 1 factory call, and all 100 received the identical object.
Awaiting a finished task returns its stored result immediately, so after initialisation get() costs one attribute check and one await on a done future — no lock, no re-run.
Verify: the factory call count is 1 and len({id(x) for x in results}) == 1.
3. Shield the shared task from caller cancellation¶
The asyncio.shield in get() is not decoration. Without it, if the first caller is cancelled — its request times out, the client disconnects — the cancellation propagates into the shared task, and every other caller waiting on it receives a CancelledError they did nothing to cause. The next caller then sees a cancelled task and fails forever.
With the shield, cancelling a caller cancels only that caller's wait; the initialisation continues for everyone else. Verified: cancelling the first caller 10 ms into a 50 ms initialisation left the factory call count at 1, and the next caller received the pool normally.
async def demo(lazy: AsyncLazy) -> None:
first = asyncio.create_task(lazy.get())
await asyncio.sleep(0.01)
first.cancel() # only this caller gives up
pool = await lazy.get() # initialisation finished anyway
The trade-off: a shielded initialisation cannot be stopped by its callers. If the factory hangs, it hangs for everyone, so put a timeout inside the factory, not around get().
Verify: cancel the first caller mid-initialisation; the factory still runs once and later callers succeed.
4. Retry after failure instead of caching it¶
As written, a failed initialisation is cached forever: the task holds the exception, and every future get() re-raises it. For a database pool that failed because the database was restarting, that turns a 5-second blip into a permanent outage of the process. Reset on failure:
async def get(self) -> T:
if self._task is None or self._failed():
self._task = asyncio.ensure_future(self._factory())
return await asyncio.shield(self._task)
def _failed(self) -> bool:
t = self._task
return t is not None and t.done() and (t.cancelled() or t.exception() is not None)
The check runs synchronously before creating a new task, so concurrent callers that arrive after a failure still converge on a single retry. What this does not do is back off: if the factory fails instantly and callers keep arriving, every caller triggers a retry. For dependencies that can be down for a while, put a circuit breaker or a minimum interval between attempts in front of the factory.
Verify: make the factory fail on its first call and succeed on its second; the first batch of callers sees the error and the next caller gets the object.
5. Close what you created¶
Lazily created resources still need deterministic cleanup, and the cleanup has to handle all three states: never started, still initialising, and ready.
async def aclose(self, close: Callable[[T], Awaitable[None]]) -> None:
task, self._task = self._task, None
if task is None:
return
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
return
if not task.cancelled() and task.exception() is None:
await close(task.result())
Wire this into the application's lifespan or AsyncExitStack so the pool is closed exactly once at shutdown, as described in managing startup and shutdown with ASGI lifespan. An initialisation still in progress at shutdown is cancelled rather than awaited, so a slow dependency cannot hold the process open.
Verify: call aclose() in each of the three states; no warnings about unclosed resources and no pending tasks remain.
Verification¶
Lazy initialisation is correct when:
- Factory calls equal 1 under any number of concurrent first callers.
- Cancelling a caller never affects other callers or the initialisation.
- A failure is retried by the next caller rather than cached for the life of the process.
- Shutdown closes the resource in every state, with no "Task was destroyed but it is pending" message.
Diagnostic Hook: count factory invocations and failures as metrics and log the initialisation duration. More than one invocation per process lifetime without a logged failure means a second code path is creating the resource outside the lazy holder; repeated failures followed by success indicate a dependency that is slow to come up at deploy time.
Pitfalls & edge cases¶
- Creating the holder at import time in a module used by several loops. The shared task belongs to the loop it was created on; awaiting it from another loop fails. Create holders per loop or per application instance.
functools.cacheon an async function. It caches the coroutine object, which can only be awaited once — covered in why functools.lru_cache breaks on coroutines.- A timeout around
get()instead of inside the factory. With shielding, the caller gives up but the hung initialisation continues and blocks every future caller. - Holding the shared task while the factory needs the holder. A factory that calls
get()on its own holder awaits itself and deadlocks.
Frequently Asked Questions¶
How do I lazily initialise a resource once in asyncio?
Store a task for the initialisation rather than its result. The first caller creates the task synchronously; every caller awaits the same task, shielded so one caller's cancellation does not affect the others. Reset the task if it fails so the next caller retries.
Why does my async lazy property create several clients?
Because the check and the assignment are separated by an await. Every caller that arrives while the first factory call is in progress sees the attribute still empty and starts its own call.
Should I use an asyncio.Lock for lazy initialisation?
A lock with a double-check works and produces one call. A shared task is simpler, needs no lock on the hot path, and lets callers be cancelled without affecting the initialisation when combined with asyncio.shield.
Why use asyncio.shield in lazy initialisation?
Without it, cancelling the first caller cancels the shared task, so every other waiting caller fails with CancelledError and the cancelled task would be cached. Shielding confines cancellation to the caller that was cancelled.
Related¶
- Coroutine Design Patterns — up to the topic overview.
- Using asyncio.shield to protect critical sections — the shield semantics this pattern relies on.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.