Bounding User-Controlled Fan-Out¶
asyncio makes fan-out effortless: gather over a list of coroutines runs them all at once. When the length of that list comes from a request — IDs in a batch lookup, recipients of a notification, URLs to preview — the client decides how much concurrent work one request creates, and an asyncio service will happily start all of it. Measured on Python 3.14 with an endpoint that fetched one item per requested ID from an upstream taking 10 ms: a single request carrying 50,000 IDs started 50,000 concurrent upstream calls, peaked at 77.7 MiB of traced memory and took 1.59 s — one request, from one client, saturating whatever was upstream. With a cap of 100 IDs per request and a shared semaphore of 20 concurrent upstream calls, the same request was rejected in 12.6 ms with no upstream calls made, and a legitimate 100-ID request completed in 73 ms at a peak of 20 concurrent calls and 0.1 MiB. This guide puts bounds on every fan-out a client can influence.
Prerequisites¶
- Python 3.11+; any async web framework.
- Bounded concurrency, from limiting concurrent requests with asyncio.Semaphore.
- The topic overview, Securing Async Services.
1. Find fan-out driven by input¶
The pattern to look for is a gather, a TaskGroup loop or as_completed whose iterable comes from the request:
@app.post("/items/batch")
async def batch(req: BatchRequest):
# one upstream call per requested id, all at once
return await asyncio.gather(*(catalog.get(item_id) for item_id in req.ids))
Measured with 50,000 IDs: 50,000 coroutines, 50,000 concurrent upstream calls at the peak, 77.7 MiB of traced memory and 1.59 s — for a single request. The upstream sees a burst equal to whatever number the client chose; connection pools, rate limits and the upstream's own capacity all fail at once, and other users' requests queue behind the flood. Authentication does not help: any logged-in user can send this. Similar shapes hide in notification fan-out (one message per recipient), link previews (one fetch per URL in a document) and GraphQL list arguments, covered in limiting GraphQL query depth and aliases.
Verify: for each endpoint, the maximum number of concurrent operations one request can create is known and finite.
2. Cap the input at the boundary¶
Reject oversized requests before any work starts. With a schema library, the cap is one declaration:
from pydantic import BaseModel, Field
MAX_IDS = 100
class BatchRequest(BaseModel):
ids: list[int] = Field(min_length=1, max_length=MAX_IDS)
@app.post("/items/batch")
async def batch(req: BatchRequest): # 422 before the handler runs if too many ids
...
Measured: the oversized request was rejected in 12.6 ms — mostly the cost of building and checking the list — with no upstream calls and 1.9 MiB of peak memory, against 1.59 s and 77.7 MiB unbounded. Choose the cap from real usage, and offer pagination or an asynchronous export for clients that legitimately need more. Validate size before content: a 50,000-element list should be rejected by its length without parsing every element, and the request body itself should be size-limited, as in limiting request body size in ASGI apps.
Verify: a request one over the cap is rejected with a 4xx and creates no tasks.
3. Bound concurrency below the cap¶
A cap limits one request; it does not limit many requests at once. Share a semaphore per upstream across all requests, so total concurrent calls stay bounded however many batches arrive:
CATALOG_SLOTS = asyncio.Semaphore(20) # per upstream, shared by all requests
async def get_item(item_id: int) -> dict:
async with CATALOG_SLOTS:
return await catalog.get(item_id)
@app.post("/items/batch")
async def batch(req: BatchRequest):
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(get_item(i)) for i in req.ids]
return [t.result() for t in tasks]
Measured: a 100-ID request completed in 73 ms with at most 20 concurrent upstream calls — five rounds of 10 ms calls plus overhead — instead of all 100 at once. Size the semaphore from what the upstream can take, divided among your service's instances. One request can still occupy all 20 slots for a while; where fairness between users matters, per-user or per-tenant limits keep one client from starving the others, as in fair scheduling across tenants in a worker pool.
Verify: under many concurrent batch requests, upstream concurrency never exceeds the semaphore size per instance.
4. Batch upstream calls instead of fanning out¶
Often the best bound is not a smaller fan-out but none at all: most upstreams that serve items by ID also serve them in bulk.
async def get_items(ids: list[int]) -> list[dict]:
results: list[dict] = []
for chunk_start in range(0, len(ids), 50):
chunk = ids[chunk_start:chunk_start + 50]
async with CATALOG_SLOTS:
results.extend(await catalog.get_many(chunk)) # one call per 50 ids
return results
A bulk call turns 100 upstream requests into 2, which reduces connection use, upstream load and the effect of per-call latency at once. Where requests from many users ask for overlapping IDs, a short-lived cache with single-flight deduplication removes repeated calls too, as in preventing cache stampedes in asyncio. The cap and semaphore still apply: bulk endpoints have their own size limits.
Verify: a 100-ID request makes ceil(100 / chunk) upstream calls, not 100.
5. Bound fan-out triggered indirectly¶
Not all fan-out is a list in a request body. Some is derived — a document whose links are previewed, a template that renders one fetch per placeholder, a user whose followers are all notified:
MAX_LINKS = 10
PREVIEW_SLOTS = asyncio.Semaphore(8)
async def render_previews(document: str) -> list[Preview]:
urls = extract_urls(document)[:MAX_LINKS] # derived input, capped too
async def one(url: str) -> Preview | None:
async with PREVIEW_SLOTS:
try:
async with asyncio.timeout(3):
return await fetch_preview(url)
except (TimeoutError, httpx.HTTPError):
return None
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(one(u)) for u in urls]
return [t.result() for t in tasks if t.result() is not None]
Derived fan-out needs the same three bounds — cap, semaphore, timeout — because the client controls it just as directly, only less visibly. Fetching URLs supplied by users also carries server-side request forgery risk: resolve and check addresses before connecting, and never let previews reach internal networks. Large fan-outs that must happen, such as notifying many followers, belong in a background job with its own pacing rather than in the request, as in Background Jobs & Task Queues.
Verify: a document with a thousand links triggers at most ten preview fetches, each bounded by a timeout.
Verification¶
Client-driven fan-out is bounded when:
- Every list from a request has a length cap, enforced before work starts.
- Every upstream has a shared semaphore sized from its capacity.
- Bulk calls replace per-item calls where the upstream supports them.
- Derived fan-out is capped and time-bounded, and very large fan-out runs as background work.
Diagnostic Hook: record, per request, the number of upstream calls it made, and alert on the maximum rather than the average. A single request making thousands of calls is the signature of an unbounded fan-out — and the average hides it completely.
Pitfalls & edge cases¶
gatherover request input. Measured: one request, 50,000 concurrent calls, 77.7 MiB.- Caps without semaphores. Many capped requests still multiply.
- Semaphores without caps. One huge request still holds memory for every pending item.
- Derived inputs. Links, recipients and placeholders are client-controlled too.
Frequently Asked Questions¶
How do I limit asyncio.gather on user input?
Cap the input length at the request boundary (for example Field(max_length=100)) and wrap each upstream call in a shared asyncio.Semaphore. In testing, the cap rejected a 50,000-item request in 12.6 ms instead of starting 50,000 concurrent calls.
Why is unbounded fan-out a security problem?
Any client — authenticated or not — chooses how much concurrent work one request creates, overwhelming upstreams, pools and memory: one 50,000-ID request used 77.7 MiB and 50,000 concurrent calls in testing.
Is a semaphore enough to protect an upstream?
It bounds concurrency but not memory: a huge request still creates a coroutine per item while it waits. Combine it with an input cap.
What about fan-out derived from content, like link previews?
Cap the derived list, use a semaphore and a per-call timeout, and check fetched addresses to avoid server-side request forgery.
Related¶
- Securing Async Services — up to the topic overview.
- Preventing regex denial of service in async handlers — CPU-side input abuse.
- Resilience, Cancellation & Error Handling — the section overview.