Securing Async Services¶
Most security advice applies to any web service. A few problems are specific to asyncio, because of the property that makes it efficient: one thread runs every request on a worker. Anything that holds that thread — a password hash, a regular expression that backtracks — becomes a denial of service for every user on the worker, and anything that makes concurrency free — gather over user input, cheap idle connections — becomes something an attacker can multiply. This section measures those problems on Python 3.14. Thirty-two concurrent bcrypt logins verified on the event loop stalled it for 5.4 s; offloaded to threads they ran at 79.7 logins per second with lag under 22 ms. A regex took 1.7 s on a 27-character input, and running it in a thread did not help — the loop still stalled 795 ms, because re holds the GIL. Eleven hundred connections that never finished their headers exhausted uvicorn's 1,024 file descriptors and every real request failed. Verifying a webhook over re-serialized JSON rejected a valid signature. A cold JWKS cache let 200 requests make 200 fetches to the identity provider. And one request with 50,000 IDs started 50,000 concurrent upstream calls.
The parent section, Resilience, Cancellation & Error Handling, covers load shedding, timeouts and bulkheads in general; this topic applies them to inputs that an adversary chooses. Each guide below reproduces one attack or mistake, measures what it does to an asyncio service, and gives the defence that held up when the attack was rerun — usually a bound on cost, count or time that rejects hostile input in milliseconds while ordinary requests continue at normal latency.
Architectural principles¶
- Nothing an attacker can trigger may run long on the event loop. CPU-heavy work moves to a pool that actually runs in parallel with the loop — and that depends on whether the code releases the GIL.
- Every client-controlled multiplier has a bound: list lengths, connection lifetimes, body sizes, concurrent operations.
- Authenticate exactly what was sent: signatures over raw bytes, compared in constant time, with replay windows.
- Cache security material carefully: single-flight refreshes, rate-limited misses, last-good fallback.
- Measure under attack conditions, not just normal load, because averages hide every one of these failures.
Execution model: why one thread changes the threat model¶
In a thread-per-request server, a slow request occupies one thread; the others keep serving. In asyncio, request handlers are coroutines sharing one thread, and they cooperate by yielding at await. A handler that does 170 ms of CPU work without yielding — one bcrypt check — delays every other request on the worker by 170 ms; thirty-two of them in a burst delayed everything by 5.4 seconds in testing. The cheapest attack on an asyncio service is therefore not volume but CPU: find an endpoint that does something expensive per request and call it a little more than the service can absorb.
Offloading is the standard answer, and it only works when the offloaded code can run while the loop runs. C extensions that release the GIL — bcrypt, argon2-cffi, most compression and hashing libraries — can be offloaded to threads. Code that holds the GIL — Python's re matching, pure-Python parsers — cannot; in a thread it still competes with the loop for the interpreter, which is why the regex measurement stalled the loop as badly from a thread as from the loop itself. Those need a process pool, or better, a rewrite that removes the cost. On the I/O side, asyncio's efficiency with idle connections means a slow-client attack does not exhaust memory or threads, but it does exhaust file descriptors, and nothing in the event loop notices until accept fails.
Pattern catalogue¶
CPU-heavy security work off the loop¶
Password hashing in a dedicated, bounded thread pool — bcrypt and argon2-cffi release the GIL — with a semaphore that rejects excess attempts. bcrypt at cost 12 cost 171 ms per check; in threads it reached 79.7 logins per second with loop lag p99 of 21 ms, and an eight-process pool kept lag at 1.1 ms. Hashing a dummy password for unknown usernames keeps response times from revealing which accounts exist. See hashing passwords without blocking the event loop.
Patterns that cannot backtrack¶
Atomic groups and possessive quantifiers (Python 3.11+), non-overlapping structure, input length caps, and a process pool with a timeout for patterns you cannot change. The possessive rewrite of ^(a+)+$ matched every test input in 0.8–2 µs, against 1.7 s for the original on 27 characters; a process pool kept the loop responsive but still spent the 772 ms of CPU in a worker, so it is a containment measure, not a fix. See preventing regex denial of service in async handlers.
Connections that cannot be held forever¶
A reverse proxy with header and body timeouts, a measured server timeout for incomplete requests, and a high, monitored descriptor limit. In testing, uvicorn's keep-alive timeout and --limit-concurrency did not close connections still sending their first headers, while Hypercorn's keep-alive timeout did and kept serving under the same attack; behaviour like this varies by server and version, so the attack itself belongs in the test suite. See defending async servers against slow clients.
Signatures over raw bytes¶
await request.body(), the provider's exact HMAC format, hmac.compare_digest, a timestamp window and event-ID deduplication. HMAC cost 1.4 µs per kilobyte, so verification belongs on the loop; the risks are verifying the wrong bytes and comparing with ==, which took measurably different times depending on where the strings differed. See verifying webhook signatures in async handlers.
A key cache that resists abuse¶
Single-flight JWKS refreshes, a minimum interval between refreshes triggered by unknown key IDs, background refresh with jitter, and last-good keys on failure. Verifying a token itself was cheap — 36 µs for RS256 — so the risk sits in the key cache: on a cold cache 200 requests made 200 fetches without single-flight and 1 with it, and 1,000 forged key IDs made 10 fetches with single-flight alone and 1 with a 30-second minimum interval. See validating JWTs with cached JWKS in asyncio.
Bounded fan-out¶
Length caps on request lists, shared per-upstream semaphores, bulk calls instead of per-item calls. The cap rejected a 50,000-ID request in 12.6 ms with no upstream calls, and the semaphore held a legitimate 100-ID request to 20 concurrent calls. See bounding user-controlled fan-out.
Testing under attack conditions¶
Each problem in this section is invisible under ordinary load tests, which send well-formed, average-sized requests at a steady rate. Testing for them means sending the inputs an attacker would. A short adversarial suite, run against staging, covers most of it: a burst of logins well above normal concurrency while measuring loop lag on the same worker; every regex-validated field filled with long near-matching strings; a few thousand connections that open and never finish their headers, with an ordinary request issued every second; webhook deliveries with modified bodies, old timestamps and repeated event IDs; tokens with random key IDs; and batch requests at, and just above, every documented limit.
What to measure is the same in every case: whether ordinary requests keep succeeding at normal latency while the adversarial traffic runs, and how quickly the hostile requests are rejected. A service that survives the suite rejects bad input in milliseconds, keeps loop lag flat, and keeps serving everyone else. The tools are the ones used throughout this section — a lag probe, a descriptor count, an upstream-call counter — and the load testing guidance applies, with one difference: here the interesting numbers are the maxima, not the percentiles.
Resource boundaries¶
- Hashing: a dedicated pool of a few workers and a semaphore of a few dozen pending attempts per instance, plus per-account and per-IP rate limits.
- Regexes: field lengths capped before matching; any validation over 10 ms treated as a bug.
- Connections: a header timeout of about 10 seconds at the proxy; a descriptor limit far above peak connections, alerted at 70%.
- Bodies: 1 MB or less on unauthenticated endpoints such as webhooks, enforced while streaming.
- Key refreshes: at most one per 30 seconds triggered by unknown key IDs; background refresh every few minutes with jitter.
- Fan-out: a documented maximum per request, and a shared semaphore per upstream sized from its capacity.
- Error detail: generic messages to clients ("invalid token", "invalid credentials") with the specific reason in logs, so rejections do not teach an attacker which check failed.
Integrated production example¶
A FastAPI application combining the defences: bounded password hashing with the same cost for unknown users, a username validator with no nested quantifiers and a length cap, a webhook endpoint that verifies a timestamped HMAC over the raw body with a size limit and deduplication, and a batch endpoint with a length cap and a shared upstream semaphore. Tested with FastAPI's test client: a valid login took 186 ms and an unknown username 173 ms; a hostile 31-character username and a 101-ID batch were rejected with 422; a 100-ID batch completed in 55 ms; the webhook accepted a signed event, reported the repeat as a duplicate and rejected a tampered body with 401; and 16 concurrent logins all succeeded in 0.40 s:
import asyncio
import hashlib
import hmac
import json
import re
import time
from concurrent.futures import ThreadPoolExecutor
import bcrypt
from fastapi import FastAPI, HTTPException, Request
from pydantic import BaseModel, Field, field_validator
app = FastAPI()
HASH_POOL = ThreadPoolExecutor(max_workers=8, thread_name_prefix="pwhash")
HASH_SLOTS = asyncio.Semaphore(32)
CATALOG_SLOTS = asyncio.Semaphore(20)
WEBHOOK_SECRET = b"whsec_example"
USERNAME = re.compile(r"^[a-z0-9](?:[a-z0-9_-]{0,30}[a-z0-9])?$") # no nested quantifiers
DUMMY_HASH = bcrypt.hashpw(b"dummy", bcrypt.gensalt(rounds=12))
class Login(BaseModel):
username: str = Field(max_length=32)
password: str = Field(max_length=256)
@field_validator("username")
@classmethod
def check_username(cls, v: str) -> str:
if not USERNAME.fullmatch(v):
raise ValueError("invalid username")
return v
class Batch(BaseModel):
ids: list[int] = Field(min_length=1, max_length=100)
@app.post("/login")
async def login(body: Login):
if HASH_SLOTS.locked():
raise HTTPException(429, "too many login attempts in progress")
stored = await users.password_hash(body.username) or DUMMY_HASH # same cost for unknown users
async with HASH_SLOTS:
loop = asyncio.get_running_loop()
ok = await loop.run_in_executor(HASH_POOL, bcrypt.checkpw, body.password.encode(), stored)
if not ok or stored is DUMMY_HASH:
raise HTTPException(401, "invalid credentials")
return {"user": body.username}
@app.post("/webhooks/payments")
async def payments(request: Request):
if int(request.headers.get("content-length", 0)) > 1 << 20:
raise HTTPException(413, "payload too large")
body = await request.body()
ts = int(request.headers.get("x-timestamp", "0"))
if abs(time.time() - ts) > 300:
raise HTTPException(400, "stale timestamp")
expected = hmac.new(WEBHOOK_SECRET, f"{ts}.".encode() + body, hashlib.sha256).hexdigest()
if not hmac.compare_digest(expected, request.headers.get("x-signature", "")):
raise HTTPException(401, "bad signature")
event = json.loads(body)
if not await processed.add_if_absent(event["id"], ttl=600):
return {"duplicate": True}
return {"accepted": event["id"]}
async def fetch_item(item_id: int) -> dict:
async with CATALOG_SLOTS:
return await catalog.get(item_id)
@app.post("/items/batch")
async def batch(body: Batch):
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(fetch_item(i)) for i in body.ids]
return [t.result() for t in tasks]
The tested version used an in-memory user table, a set for deduplication and a 10 ms sleep for the catalog; production replaces them with the user store, Redis SET NX and the real client, without changing the structure. JWT verification would sit in a shared dependency as shown in the JWKS guide, and the slow-client defence lives in front of the application, in the proxy.
Diagnostic hook callout¶
The metrics that reveal these attacks are maxima and counts, not averages:
- Maximum event-loop lag per minute — a single multi-second stall points at CPU work on the loop (a hash, a regex) and the request that triggered it.
- Open file descriptors against the limit — rising without a matching rise in completed requests means slow clients or leaks.
- Maximum upstream calls per request — any request far above the documented cap means an unbounded fan-out path.
- Rejections by reason — bad signatures, stale timestamps, unknown key IDs, oversized inputs; a sudden rise in one is either an attack or a regression.
Alert on loop lag above 250 ms, descriptors above 70% of the limit, any request exceeding its fan-out cap, and JWKS fetches above a handful per hour per instance.
Failure modes¶
| Failure mode | Root cause | Detection | Fix |
|---|---|---|---|
| Whole worker stalls during logins | Hashing on the event loop | Max loop lag spikes with logins | Dedicated, bounded hash pool |
| Stall persists after moving to a thread | Code holds the GIL (re) |
Loop lag unchanged from a thread | Rewrite the pattern; process pool |
| Real requests fail, CPU idle | Descriptors held by slow clients | fds near limit, no completed requests | Proxy header timeouts; higher limit |
| Valid webhooks rejected | HMAC over parsed/re-serialized data | Signature failures after a change | Verify raw bytes |
| Identity provider flooded | Uncoordinated or unbounded JWKS refreshes | Fetches per instance | Single-flight + minimum interval |
| One request floods an upstream | Fan-out sized by input | Upstream calls per request | Caps, semaphores, bulk calls |
Frequently Asked Questions¶
What security issues are specific to asyncio services?
Anything that holds the single event-loop thread — password hashing, backtracking regexes — stalls every request on the worker, and cheap concurrency lets clients multiply work through fan-out or held connections. In testing, 32 concurrent bcrypt logins on the loop stalled it for 5.4 s.
Is moving CPU work to a thread always enough in asyncio?
No. It works for C code that releases the GIL, such as bcrypt; Python's re holds the GIL, and a slow regex in a thread still stalled the loop for 795 ms in testing. Use a process pool or rewrite the work.
Are asyncio servers immune to slowloris?
They handle idle connections cheaply, but each holds a file descriptor: 1,100 slow clients exhausted a 1,024 limit in uvicorn and every real request failed. Put a proxy with header timeouts in front.
How do I verify webhooks in an async framework?
Compute the HMAC over the raw request bytes, compare with hmac.compare_digest, reject stale timestamps and deduplicate event IDs; verifying re-serialized JSON rejected a valid signature in testing.
How do I stop one request from overwhelming an upstream?
Cap list lengths at the request boundary and share a semaphore per upstream across requests: a 50,000-ID request was rejected in 12.6 ms instead of starting 50,000 concurrent calls.
Related¶
- Hashing passwords without blocking the event loop — CPU-heavy crypto off the loop.
- Preventing regex denial of service in async handlers — patterns that cannot explode.
- Defending async servers against slow clients — descriptors and timeouts.
- Verifying webhook signatures in async handlers — raw bytes, constant time.
- Validating JWTs with cached JWKS in asyncio — a key cache attackers cannot drive.
- Bounding user-controlled fan-out — caps and semaphores.
- Circuit Breakers & Bulkheads — containment when defences are not enough.
- Resilience, Cancellation & Error Handling — the parent section.