Skip to content

Why functools.lru_cache Breaks on Coroutines

Decorating an async def with functools.lru_cache looks like it should memoise the result. It memoises the coroutine object — the thing get_user(1) returns before anything runs — and a coroutine can only be awaited once. The first await get_user(1) works; the second raises RuntimeError: cannot reuse already awaited coroutine. Worse, some call patterns hide the bug: asyncio.gather(get_user(2), get_user(2)) received the same cached coroutine twice, deduplicated it internally, and succeeded — so a test written that way passes while production fails on the first repeat request. The fix is to cache a task, which can be awaited any number of times. A 20-line decorator doing that served 50 concurrent first callers with one underlying call, and retried after failures instead of caching them; the async-lru package's alru_cache also made one call for 50 concurrent callers.

Prerequisites

1. See exactly what lru_cache stores

import asyncio
import functools


@functools.lru_cache(maxsize=128)
async def get_user(uid: int) -> dict:
    await asyncio.sleep(0.01)
    return {"id": uid}


async def main() -> None:
    print(await get_user(1))          # {'id': 1}
    print(await get_user(1))          # RuntimeError: cannot reuse already awaited coroutine


asyncio.run(main())

lru_cache wraps the function call. Calling an async def function does no work; it creates a coroutine object, and that object is what gets cached under the key (1,). The second call returns the same, already-exhausted coroutine. The same applies to functools.cache and functools.cached_property on async methods.

Because asyncio.gather deduplicates identical awaitables, gather(get_user(2), get_user(2)) wraps the one cached coroutine in one task and hands the result to both slots — it succeeded in the test. A unit test of "concurrent callers share a result" can therefore pass against a decorator that fails on the next sequential call.

Verify: call the decorated function twice sequentially in a test; it must raise, and that test should be in your suite as a guard against reintroducing the decorator.

What lru_cache actually caches for an async def A sequence of 6 messages between 3 participants. What lru_cache actually caches for an async def caller lru_cache coroutine object get_user(1) create coroutine, cache it await: runs, returns result get_user(1) again same coroutine from cache await: RuntimeError, already awaited The cache key maps to an object that can only be consumed once.

2. Cache a task instead

A task wraps the coroutine, runs it once, and stores the result; awaiting a finished task returns that result immediately, as often as you like. Cache the task:

import asyncio
import functools


def async_cache(fn):
    cache: dict[tuple, asyncio.Task] = {}

    @functools.wraps(fn)
    async def wrapper(*args):
        task = cache.get(args)
        if task is None:
            task = cache[args] = asyncio.ensure_future(fn(*args))

            def drop_on_error(t: asyncio.Task) -> None:
                if t.cancelled() or t.exception() is not None:
                    cache.pop(args, None)          # never cache failures

            task.add_done_callback(drop_on_error)
        return await asyncio.shield(task)

    wrapper.cache = cache
    return wrapper

Measured: 50 concurrent first calls for one key made 1 underlying call — later callers find the in-flight task and wait on it, which is single-flight deduplication for free. A key whose call failed twice was retried each time and cached on the first success, three calls in total, because the done-callback evicts failed tasks.

The shield stops one caller's cancellation from cancelling the shared task for every other caller; without it, a client disconnecting during the first fetch would fail every concurrent request for the same key. The same reasoning appears in lazy async initialization with a shared task.

Verify: sequential repeats return the cached value; concurrent first calls produce one underlying call; a failure is not cached.

3. Bound the cache and add expiry

The decorator above grows forever and never refreshes. Production caches need a size limit and usually a TTL:

import time
from collections import OrderedDict


def async_ttl_cache(maxsize: int = 1024, ttl: float = 60.0):
    def decorator(fn):
        cache: OrderedDict[tuple, tuple[float, asyncio.Task]] = OrderedDict()

        @functools.wraps(fn)
        async def wrapper(*args):
            now = time.monotonic()
            entry = cache.get(args)
            if entry is not None and entry[0] > now:
                cache.move_to_end(args)                     # LRU touch
                return await asyncio.shield(entry[1])
            task = asyncio.ensure_future(fn(*args))
            cache[args] = (now + ttl, task)
            cache.move_to_end(args)
            while len(cache) > maxsize:
                cache.popitem(last=False)                   # evict least recently used
            task.add_done_callback(
                lambda t: (t.cancelled() or t.exception()) and cache.pop(args, None))
            return await asyncio.shield(task)

        return wrapper
    return decorator

Evicting an entry whose task is still running is safe: callers already awaiting it keep their reference, and a new caller simply starts a fresh call. The full treatment — including what happens when the cache is shared across workers — is in building an async TTL cache decorator.

Verify: after ttl seconds a repeat call triggers a fresh underlying call, and the cache never exceeds maxsize entries.

Ways to cache an async function A grid of 4 rows by 5 columns. Ways to cache an async function approach repeat await concurrent callers failures expiry functools.lru_cache RuntimeError dedup by accident cached no cache the task works one call evicted no task cache + TTL + LRU works one call evicted yes async_lru.alru_cache works one call (measured) per library docs ttl option Every working option caches something awaitable many times — a task or a stored result — never a coroutine.

4. Or use async-lru

If you would rather not maintain the decorator, async-lru provides alru_cache with the familiar lru_cache interface plus a ttl option:

# pip install async-lru
from async_lru import alru_cache


@alru_cache(maxsize=1024, ttl=60)
async def get_user(uid: int) -> dict:
    return await db.fetch_user(uid)


await get_user.cache_close()      # on shutdown: cancels in-flight calls cleanly

Measured with async-lru 2.3.0: 50 concurrent first calls produced 1 underlying call. Read its documentation for its exact semantics on exceptions and cancellation before relying on them, and wire cache_close() into your shutdown path. A library decorator is the right default for function-level caching; the hand-written version earns its place when you need custom keys, metrics, or the shield behaviour spelled out.

Verify: the library's behaviour on a failing call matches what your service needs — cached, retried, or neither.

5. Watch for methods and unhashable arguments

Two details catch people moving from lru_cache to an async cache:

  • Methods. Caching an instance method keys on self, which keeps every instance alive as long as the cache holds it — a leak if instances are per-request. Cache a module-level function keyed on the identifying fields instead, or hold a cache per instance.
  • Unhashable arguments. Dicts and lists cannot be keys. Normalise them into a tuple or a stable string before calling the cached function, rather than making the cache stringify arbitrary objects.
class UserService:
    def __init__(self, db) -> None:
        self._db = db
        self.get_user = async_ttl_cache(maxsize=10_000, ttl=30)(self._get_user)   # per instance

    async def _get_user(self, uid: int) -> dict:
        return await self._db.fetch_user(uid)

Binding the cache per instance makes the cache's lifetime equal the service's lifetime, which is exactly what you want for objects created once at startup.

Verify: creating and discarding many service instances does not grow memory; a single long-lived instance caches across calls.

Which caching approach for this coroutine? A decision on What do you need from the cache with 3 outcomes. Which caching approach for this coroutine? What do you need from the cache? plain memoisation async_lru.alru_cache maxsize and ttl custom keys, metrics, shield your own task cache 20-40 lines shared across processes Redis-backed cache plus local tier Never decorate an async def with functools caches; they store the wrong object.

Verification

Caching is correct when:

  • No functools.lru_cache, cache or cached_property decorates an async def.
  • Sequential repeats return cached results without error.
  • Concurrent first callers produce one underlying call.
  • Failures are not cached, and the cache is bounded in size.

Diagnostic Hook: export hits, misses, in-flight joins (callers who found a running task) and evictions per cached function. A high in-flight-join count means the cache is doing single-flight work under concurrency; a miss rate that climbs after a deploy usually means the cache key changed shape — an extra argument, or a new object type that no longer compares equal.

Pitfalls & edge cases

  • Tests that only use gather with identical calls. They pass even with lru_cache, because gather deduplicates the cached coroutine.
  • Caching exceptions. A transient failure becomes permanent for the cache's lifetime.
  • No shield. One caller's cancellation fails every caller sharing the in-flight task.
  • Caching across event loops. Cached tasks belong to the loop that created them; clear the cache between asyncio.run() calls in tests.

Frequently Asked Questions

Can I use functools.lru_cache on an async function?

No. It caches the coroutine object returned by the call, and a coroutine can only be awaited once, so the second await raises RuntimeError: cannot reuse already awaited coroutine. Cache a task or the awaited result instead.

What is the async equivalent of lru_cache?

The async-lru package's alru_cache decorator, which supports maxsize and ttl and deduplicates concurrent calls. You can also write a small decorator that caches an asyncio task per key.

Why does caching a task work when caching a coroutine does not?

A task runs its coroutine once and stores the result; awaiting a finished task returns the stored result every time. A coroutine object is consumed by its first await.

Should an async cache store exceptions?

Usually not. Evict the entry when the task fails or is cancelled, so the next caller retries, otherwise a transient error is served from the cache until the entry expires.