Skip to content

Scoping GraphQL DataLoaders per Request

A DataLoader has two jobs: batching loads made in the same tick, and caching every key it has loaded so a request never fetches the same record twice. The cache is what makes its scope matter. Created per request, the cache lives for one request and then disappears. Created once at module level — a tempting "optimization" — it becomes an unbounded, never-invalidated cache shared by every user. Measured with Strawberry 0.329 on Python 3.14: with a module-level loader, a user record renamed between two requests was still returned with its old name by the second request; with a loader created in the request context, the second request saw the new name. After 100 requests that each loaded 1,000 different users, the shared loader held 100,000 cached entries and the process retained 23.8 MiB more memory than with per-request loaders, which held nothing between requests. This guide scopes loaders correctly and adds deliberate caching where it is wanted.

Prerequisites

1. See the stale read

A loader created at import time looks harmless — it is "just a batching helper":

from strawberry.dataloader import DataLoader

users_loader = DataLoader(load_fn=load_users)       # module level: shared by all requests


@strawberry.type
class Query:
    @strawberry.field
    async def user(self, id: int) -> User:
        row = await users_loader.load(id)
        return User(id=id, name=row["name"], email=row["email"])

Measured: the first request loaded user 7 as user7; the record was then renamed to renamed; the second request, through the same loader, still returned user7. The loader's cache maps each key to the future from its first load, and nothing ever evicts it. In a real service the update might be the user changing their own profile, and the stale read appears to them as a save that did not work — until the process restarts. With the loader created in the context getter instead, the second request read renamed.

Verify: update a record through a mutation, then query it in a new request; the response shows the new value.

Module-level versus per-request DataLoader A grid of 4 rows by 3 columns. Module-level versus per-request DataLoader behaviour module-level loader per-request loader second request after a rename old name (user7) new name (renamed) cache entries after 100 x 1,000 users 100,000 0 between requests extra retained memory +23.8 MiB baseline data visible across users yes, anything any request loaded only within one request Strawberry 0.329, Python 3.14; memory measured with tracemalloc relative to the per-request run.

2. Create loaders in the context getter

The fix is to construct loaders where the request is constructed:

from functools import partial


async def get_context(request: Request) -> dict:
    pool = request.app.state.pool
    return {
        "request": request,
        "user": await authenticate(request),
        "loaders": Loaders(pool),
    }


class Loaders:
    """Every DataLoader for one request; discarded with it."""
    def __init__(self, pool) -> None:
        self.users = DataLoader(load_fn=partial(load_users, pool))
        self.posts_by_author = DataLoader(load_fn=partial(load_posts_by_author, pool))


@strawberry.type
class Query:
    @strawberry.field
    async def user(self, info: strawberry.Info, id: int) -> User:
        return await info.context["loaders"].users.load(id)

Constructing a DataLoader is cheap — measured at about 0.26 µs each — so creating several per request costs nothing measurable next to executing the query. Grouping them in one object keeps resolvers from constructing their own (which would defeat batching) and makes it obvious, in review, where loaders come from. Within the request the cache still deduplicates: a thousand posts by twenty authors load twenty authors once each.

Verify: no DataLoader( call appears outside the context-building code.

3. Keep authorization out of the shared path

A shared loader is also a data-exposure risk. Loaders cache by key, not by who asked; if an earlier request loaded a record with permissions the current user lacks, a shared cache hands it over:

# Wrong: the cache key is just the ID
async def load_documents(ids: list[int]) -> list[Document]:
    return await db.fetch_documents(ids)                   # no notion of the viewer


# Right: the loader is per request, and authorization is applied per viewer
class Loaders:
    def __init__(self, pool, viewer: User) -> None:
        self.documents = DataLoader(load_fn=partial(load_documents_for, pool, viewer.id))


async def load_documents_for(pool, viewer_id: int, ids: list[int]) -> list[Document | None]:
    rows = await pool.fetch(
        "SELECT d.* FROM document d JOIN acl a ON a.document_id = d.id "
        "WHERE d.id = ANY($1::int[]) AND a.user_id = $2", ids, viewer_id)
    found = {r["id"]: Document(**r) for r in rows}
    return [found.get(i) for i in ids]                     # None for forbidden or missing

With per-request loaders bound to the viewer, every load is filtered for the user making the request, and nothing loaded for one user can be served to another. The measured stale-name case in step 1 is the same mechanism as a cross-user leak — a value cached under one request's conditions served under another's.

Verify: a test where an administrator's request runs before a regular user's shows the regular user only what they may see.

One request's DataLoader lifecycle A flow of 4 stages. One request's DataLoader lifecycle context getter authenticate, build Loaders(viewer) resolvers await loaders.x.load(key) batch + per-request cache one query, viewer-filtered response sent context and caches discarded The cache lives exactly as long as the request that filled it.

4. Prime and clear inside mutations

Within one request, a mutation that changes a record should update the loader's view of it, or a later field in the same response reads the old value:

@strawberry.type
class Mutation:
    @strawberry.mutation
    async def rename_user(self, info: strawberry.Info, id: int, name: str) -> User:
        row = await info.context["db"].fetchrow(
            "UPDATE gql_user SET name = $2 WHERE id = $1 RETURNING id, name, email", id, name)
        loaders = info.context["loaders"]
        loaders.users.clear(id)                    # forget any earlier load in this request
        loaders.users.prime(id, dict(row))         # and seed the fresh value
        return User(**row)

clear removes a key, clear_all empties the cache, and prime inserts a value without a load. GraphQL executes the fields of a mutation operation serially, so a mutation followed by a query field in the same document reads whatever the loader holds at that point — which is why the clear-then-prime belongs in the mutation itself.

Verify: a single document that renames a user and then selects that user's name returns the new name.

5. Add a deliberate cross-request cache where it pays

Some data genuinely should be cached across requests: reference data that rarely changes, expensive computations, public catalogue entries. Put that cache under the per-request loader, with an explicit TTL and size bound:

from cachetools import TTLCache

CATALOGUE_CACHE: TTLCache[int, dict] = TTLCache(maxsize=50_000, ttl=60)


async def load_products(pool, ids: list[int]) -> list[dict | None]:
    missing = [i for i in ids if i not in CATALOGUE_CACHE]
    if missing:
        rows = await pool.fetch("SELECT * FROM product WHERE id = ANY($1::int[])", missing)
        for r in rows:
            CATALOGUE_CACHE[r["id"]] = dict(r)
    return [CATALOGUE_CACHE.get(i) for i in ids]

The per-request loader still batches and deduplicates; the shared cache below it has a bound on memory, a bound on staleness, and contains only data that is the same for every viewer. Invalidating it across worker processes needs a signal, as in invalidating caches across async workers, and concurrent misses for the same key need stampede protection, as in preventing cache stampedes in asyncio.

Verify: the shared cache's size stays below its bound under load, and a changed product is visible within the TTL.

Where should this data be cached? A decision on Who may see the cached value, and for how long with 4 outcomes. Where should this data be cached? Who may see the cached value, and for how long? this request only per-request DataLoader cache automatic depends on the viewer never across requests per-request only same for everyone, changes rarely bounded TTL cache under the loader explicit staleness just changed by a mutation clear() then prime() same request sees it A DataLoader is a batching tool with a request-sized cache, not an application cache.

Verification

DataLoaders are scoped correctly when:

  • Every loader is created in the context getter, never at module level.
  • Loaders for access-controlled data are bound to the viewer.
  • Mutations clear and prime the affected keys.
  • Cross-request caching is explicit, bounded in size and time, and limited to viewer-independent data.

Diagnostic Hook: track the process's RSS over a day with steady traffic. A module-level loader shows up as steady growth that resets only on restart; per-request loaders keep memory flat. Pair it with a synthetic check that updates a record and immediately reads it back in a new request.

Pitfalls & edge cases

  • Module-level loaders. Measured: stale reads after updates and 100,000 retained entries.
  • Loaders created inside resolvers. Each resolver gets its own, so nothing batches.
  • Caching viewer-specific data globally. It can serve one user's data to another.
  • Unbounded "temporary" caches. Every shared cache needs a size and a TTL.

Frequently Asked Questions

Should a GraphQL DataLoader be global or per request?

Per request. A module-level loader returned a stale name after the record was updated and kept 100,000 cached users after 100 requests; a loader created in the context getter saw the update and held nothing between requests.

Why does my GraphQL API return old data after a mutation?

Usually a DataLoader whose cache outlives the request, or a mutation that did not clear and prime the loader. Create loaders per request and call loader.clear(key) and loader.prime(key, value) in mutations.

Can a shared DataLoader leak data between users?

Yes: it caches by key, not by viewer, so a record loaded under one user's permissions can be served to another. Bind per-request loaders to the viewer and filter in the batch function.

How do I cache GraphQL data across requests safely?

Put a bounded TTL cache beneath the per-request DataLoader, only for data that is the same for every viewer, and invalidate it across workers when the data changes.