Async GraphQL in Python¶
GraphQL moves query planning from the server to the client: the client sends a document describing the data it wants, and the server's executor walks it, calling a resolver for every field. On asyncio that executor is concurrent — sibling async resolvers run together — which is a strength for latency and a liability for load, because one document can start a very large amount of concurrent work on one event loop. This section measures the moving parts with Strawberry 0.329 on Python 3.14, FastAPI and uvicorn, with asyncpg and PostgreSQL 17 behind it. A resolver that blocked for 50 ms held the server to 19 requests per second where async resolvers reached 277–327. A query for 1,000 users and their posts issued 1,001 SQL queries without a DataLoader and 2 with one, and at a simulated 5 ms database round trip took 619 ms against 128 ms. Each extra level of a recursive field multiplied work by ten — 73 ms to 947 ms — until a depth limit rejected the query in about 1 ms. A module-level DataLoader returned stale data after an update. Strawberry's executor cost about 2.4–3.0 µs per field before any I/O, and its OpenTelemetry extension raised that to 14 µs. Subscriptions delivered events to 500 clients at an 11.6 ms p50 and cleaned up all of them on disconnect.
The parent section, Network I/O & Protocol Handling, covers the ASGI servers, WebSockets and database drivers GraphQL sits on; this topic is about the GraphQL layer itself.
Architectural principles¶
- Every resolver that waits is async. A
defresolver runs on the event loop; if it blocks, every request on the worker waits. - Every keyed lookup goes through a per-request DataLoader, created in the context getter and discarded with the request.
- Every operation is bounded — depth, aliases, tokens, page sizes — and runs under a deadline.
- Instrumentation follows the I/O: time awaitable resolvers and loader batches, not every field.
- Subscriptions are async generators that register in
tryand clean up infinally, fed by publishers that never await subscribers.
Execution model: concurrent resolvers on one loop¶
Strawberry parses and validates the document, then executes it field by field. For each field it calls the resolver; if the resolver returns an awaitable, the executor gathers it with its siblings and awaits them together. Three aliased fields that each awaited 50 ms took 66 ms at p50, not 150. Nested fields wait for their parent, so the critical path through a document is its depth, while its width runs concurrently. Plain def resolvers are called inline on the loop thread — perfect for attribute access, and the source of the 19 requests per second when one of them slept.
That concurrency is what makes the N+1 problem look tolerable in development: 1,001 per-user queries run in parallel across the connection pool, so a local test takes 180–250 ms rather than seconds. The database still parses and plans a thousand statements, and as round-trip time grows the gap widens — at 5 ms it was 619 ms against 128 ms for the batched version. DataLoader turns the concurrency into an advantage: because all posts resolvers start in the same tick, all their load(id) calls land in one batch. Underneath everything is the executor's own cost: resolving 32,000 in-memory fields took 78 ms with no I/O at all, about 2.4–2.5 µs per field, and closer to 3 µs with async list resolvers. Large responses are slow because they are large.
Pattern catalogue¶
Async resolvers on FastAPI¶
@strawberry.type
class Query:
@strawberry.field
async def product(self, info: strawberry.Info, id: int) -> Product | None:
row = await info.context["pool"].fetchrow("SELECT ... WHERE id = $1", id)
return Product(**row) if row else None
Blocking libraries run through asyncio.to_thread from an async resolver, which raised throughput from 19 to 327 requests per second for a 50 ms blocking call. Per-request state — the viewer, the pool, the loaders — comes from the context getter, so resolvers never reach for module-level state. See building async GraphQL APIs with Strawberry.
DataLoader batching¶
@strawberry.field
async def posts(self, info: strawberry.Info) -> list[Post]:
return await info.context["loaders"].posts_by_author.load(self.id)
One query per nesting level regardless of list size, with results returned in key order. See batching GraphQL resolvers with DataLoader.
Per-request loader scope¶
Loaders built in the context getter, bound to the viewer, cleared and primed by mutations. A loader's cache maps each key to the result of its first load for as long as the loader lives, so scope is a correctness decision: per request, it deduplicates within one response; at module level, it served a renamed user's old name to the next request and kept 100,000 entries after 100 requests. Data that genuinely should be shared across requests belongs in an explicit, bounded TTL cache beneath the loader. See scoping DataLoaders per request.
Operation limits and deadlines¶
schema = strawberry.Schema(query=Query, extensions=[
lambda: QueryDepthLimiter(max_depth=6),
lambda: MaxAliasesLimiter(max_alias_count=15),
lambda: MaxTokensLimiter(max_token_count=2000),
])
Plus clamped page sizes and an execution deadline returned as a GraphQL error. See limiting GraphQL query depth and aliases.
Subscriptions¶
Async generators with bounded per-subscriber queues and cleanup in finally. The publisher pushes with put_nowait and never awaits a subscriber, so a stalled client fills only its own queue and loses only its own oldest events; authentication happens once in the WebSocket connection_init, and per-user caps bound how many generators one client can hold open. See streaming GraphQL subscriptions over WebSockets.
Resolver timing and tracing¶
Time only awaitable resolvers; trace loader batches, not individual resolver calls. Skipping plain attribute fields halved the timing overhead to about 0.4 µs per field, and a span per loader batch keeps a trace's size proportional to nesting depth rather than to the number of items, which matters when a single response contains thousands of resolver calls. See timing and tracing async GraphQL resolvers.
Designing the schema for async execution¶
Most of the failures measured in this section are schema decisions as much as code decisions, and they are cheaper to get right before clients depend on the shape. Four habits keep a schema friendly to an async executor.
Make lists connections, with a bounded page size. Every list field that can grow takes first/after arguments with a server-side maximum. Bare list[User] fields are fine only for collections that are small by nature — a user's roles, an order's line items — and should be documented as such.
Model recursion explicitly. Relationships such as friends, replies or child categories are where depth becomes exponential. Give them a smaller default page than top-level lists, and set the depth limit from the deepest query a real client needs rather than from a round number.
Put expensive computation behind its own field. A field that aggregates, ranks or calls a slow service should be separate from the cheap fields of the same type, so clients that do not need it do not pay for it, and so it can carry its own timeout and nullability.
Decide nullability from failure, not from data. A field that depends on an optional service should be nullable so it can degrade to null with an error entry; a field that must never be silently missing should be non-null so its failure is loud. Nullability is how a GraphQL schema expresses which dependencies are allowed to fail.
Bulk work deserves a mention of its own. GraphQL is excellent for screens — many small, related pieces of data in one round trip — and poor for exports: at about 3 µs of executor time per field, a 100,000-row export with ten columns spends three seconds in field resolution before serialization. For those, a plain streaming endpoint that writes rows directly, as in streaming responses with Starlette and FastAPI, is both faster and easier to bound than any GraphQL query.
Resource boundaries¶
- Fields per response: at roughly 3 µs each on one event loop, 30,000 fields is about 100 ms of CPU before any I/O. Page sizes and depth limits bound it.
- Concurrent upstream calls per request: sibling async fields overlap, so one request can start as many calls as it has siblings; aliases multiply them. Alias limits and a shared semaphore bound it.
- Database statements per request: one per nesting level with DataLoaders; anything that scales with list size is a missing loader.
- Connection pool: concurrent loader batches and resolvers share it; size it for concurrent requests times nesting depth, not for items.
- Subscriptions: each holds a WebSocket, a generator and a queue; cap them per user and keep queues bounded.
- Execution time: a deadline per operation, below the server's and proxy's timeouts.
Integrated production example¶
A FastAPI application combining the patterns: per-request loaders built in the context getter, clamped page sizes, depth, alias and token limits as extension factories, a deadline returned as a GraphQL error, and timing of async resolvers only. In a test against PostgreSQL it returned 1,000 users with their 10,000 posts in 153.8 ms with two SQL statements, clamped first: 100000 to 1,000 rows, rejected a 20-alias document in 3.8 ms, and served a 100-user query (about 3,200 fields) at 53–57 requests per second on one worker at 20 concurrent:
import asyncio
import inspect
import time
from collections import defaultdict
from contextlib import asynccontextmanager
from functools import partial
import asyncpg
import strawberry
from fastapi import FastAPI, Request
from graphql import GraphQLError
from strawberry.dataloader import DataLoader
from strawberry.extensions import MaxAliasesLimiter, MaxTokensLimiter, QueryDepthLimiter, SchemaExtension
from strawberry.fastapi import GraphQLRouter
from strawberry.types import ExecutionResult
FIELD_SECONDS: dict[str, list[float]] = defaultdict(list)
async def load_posts(pool, author_ids: list[int]) -> list[list["Post"]]:
rows = await pool.fetch("SELECT id, author_id, title, likes FROM gql_post "
"WHERE author_id = ANY($1::int[])", list(author_ids))
by_author: dict[int, list[Post]] = defaultdict(list)
for r in rows:
by_author[r["author_id"]].append(Post(id=r["id"], title=r["title"], likes=r["likes"]))
return [by_author[a] for a in author_ids]
class Loaders: # one set per request
def __init__(self, pool) -> None:
self.posts_by_author = DataLoader(load_fn=partial(load_posts, pool), max_batch_size=1000)
@strawberry.type
class Post:
id: int
title: str
likes: int
@strawberry.type
class User:
id: int
name: str
@strawberry.field
async def posts(self, info: strawberry.Info) -> list[Post]:
return await info.context["loaders"].posts_by_author.load(self.id)
@strawberry.type
class Query:
@strawberry.field
async def users(self, info: strawberry.Info, first: int = 20) -> list[User]:
first = max(1, min(first, 1000)) # clamp page size
rows = await info.context["pool"].fetch("SELECT id, name FROM gql_user ORDER BY id LIMIT $1", first)
return [User(id=r["id"], name=r["name"]) for r in rows]
class AsyncResolverTimer(SchemaExtension):
def resolve(self, _next, root, info, *args, **kwargs):
result = _next(root, info, *args, **kwargs)
if not inspect.isawaitable(result):
return result
start, key = time.perf_counter(), f"{info.parent_type.name}.{info.field_name}"
async def timed():
try:
return await result
finally:
FIELD_SECONDS[key].append(time.perf_counter() - start)
return timed()
schema = strawberry.Schema(query=Query, extensions=[
lambda: QueryDepthLimiter(max_depth=6),
lambda: MaxAliasesLimiter(max_alias_count=15),
lambda: MaxTokensLimiter(max_token_count=2000),
AsyncResolverTimer,
])
class TimeoutRouter(GraphQLRouter):
async def execute_operation(self, *args, **kwargs):
try:
async with asyncio.timeout(5.0):
return await super().execute_operation(*args, **kwargs)
except TimeoutError:
return ExecutionResult(data=None, errors=[GraphQLError("query exceeded its time budget")])
@asynccontextmanager
async def lifespan(app: FastAPI):
app.state.pool = await asyncpg.create_pool(DSN, min_size=5, max_size=20)
yield
await app.state.pool.close()
async def get_context(request: Request) -> dict:
pool = request.app.state.pool
return {"request": request, "pool": pool, "loaders": Loaders(pool)}
app = FastAPI(lifespan=lifespan)
app.include_router(TimeoutRouter(schema, context_getter=get_context), prefix="/graphql")
The field timer recorded two Query.users calls and 1,000 User.posts calls for the first request, with a p50 of about 75 ms for posts — mostly time spent waiting for the batch to be dispatched and the executor to reach each user, which is exactly the kind of figure that tells you the query, not the database, is the cost. At 53–57 requests per second for a 3,200-field response, one worker spent roughly 18 ms of CPU per request: the next lever for this endpoint is a smaller default page, not a faster query.
Diagnostic hook callout¶
Per operation name, record: SQL statements, resolved fields, wall time, and rejections by limit. Per async resolver, record a latency histogram. Per process, record event-loop lag and active subscriptions. Alert on:
- SQL statements that scale with response size — a missing or bypassed DataLoader.
- Event-loop lag rising with traffic — a blocking
defresolver, or responses large enough that execution itself saturates the loop. - Rejections concentrated on one client — an abusive or misconfigured client hitting depth or alias limits.
- Active subscriptions above connected clients — generators that never clean up.
- Deadline errors above a small baseline — operations whose data-dependent cost exceeds the budget.
Failure modes¶
| Failure mode | Root cause | Detection | Fix |
|---|---|---|---|
| Throughput collapses under load | Blocking call in a def resolver |
Event-loop lag; 19 vs 277 req/s | async def + async client or to_thread |
| Database overloaded by one query type | N+1 resolvers | Statements per request track list size | Per-request DataLoader |
| Stale or foreign data | Module-level DataLoader cache | Read-after-write check fails | Loaders in the context getter, bound to the viewer |
| One query stalls a worker | Unbounded depth, aliases or page size | Resolver counts per operation | Limits as extension factories; clamped pages |
| Long requests with fast resolvers | Executor cost on huge responses | Fields per response × ~3 µs | Pagination, smaller defaults |
| Memory grows with subscriptions | Cleanup outside finally; unbounded queues |
Subscriptions above connections | try/finally; bounded queues |
| Tracing slows large queries | Span per resolver call | Latency with and without the extension | Time async resolvers; span per loader batch |
Frequently Asked Questions¶
Which async GraphQL library should I use in Python?
This section uses Strawberry, which builds schemas from type hints, runs on asyncio, integrates with FastAPI and supports subscriptions; Ariadne and Graphene-based stacks follow the same execution model, and the same rules about blocking resolvers and DataLoaders apply.
Does async GraphQL solve the N+1 problem?
No. Async resolvers run the N+1 queries concurrently, which hides some latency, but a 1,000-user query still issued 1,001 SQL statements; a DataLoader reduced that to 2.
How do I stop expensive GraphQL queries?
Add QueryDepthLimiter, MaxAliasesLimiter and MaxTokensLimiter as extension factories, clamp list arguments, and run each operation under a deadline. A depth-5 recursive query that took 947 ms was rejected in about 1 ms.
How fast is Strawberry GraphQL?
Its executor cost about 2.4–3.0 µs per resolved field in testing, so a 32,000-field response took 78–96 ms before any I/O; a trivial query served about 1,040 requests per second on one worker.
Should DataLoaders be shared between requests?
No. A module-level loader served stale data after an update and cached 100,000 users after 100 requests. Create loaders per request.
Related¶
- Building async GraphQL APIs with Strawberry — resolvers that keep the loop free.
- Batching GraphQL resolvers with DataLoader — N+1 to one query per level.
- Scoping DataLoaders per request — no stale or foreign data.
- Limiting GraphQL query depth and aliases — bounding every operation.
- Streaming GraphQL subscriptions over WebSockets — fan-out and cleanup.
- Timing and tracing async GraphQL resolvers — affordable visibility.
- ASGI Servers & Frameworks — the server layer underneath.
- Network I/O & Protocol Handling — the parent section.