Skip to content

Securing Async Services

Most security advice applies to any web service. A few problems are specific to asyncio, because of the property that makes it efficient: one thread runs every request on a worker. Anything that holds that thread — a password hash, a regular expression that backtracks — becomes a denial of service for every user on the worker, and anything that makes concurrency free — gather over user input, cheap idle connections — becomes something an attacker can multiply. This section measures those problems on Python 3.14. Thirty-two concurrent bcrypt logins verified on the event loop stalled it for 5.4 s; offloaded to threads they ran at 79.7 logins per second with lag under 22 ms. A regex took 1.7 s on a 27-character input, and running it in a thread did not help — the loop still stalled 795 ms, because re holds the GIL. Eleven hundred connections that never finished their headers exhausted uvicorn's 1,024 file descriptors and every real request failed. Verifying a webhook over re-serialized JSON rejected a valid signature. A cold JWKS cache let 200 requests make 200 fetches to the identity provider. And one request with 50,000 IDs started 50,000 concurrent upstream calls.

The parent section, Resilience, Cancellation & Error Handling, covers load shedding, timeouts and bulkheads in general; this topic applies them to inputs that an adversary chooses. Each guide below reproduces one attack or mistake, measures what it does to an asyncio service, and gives the defence that held up when the attack was rerun — usually a bound on cost, count or time that rejects hostile input in milliseconds while ordinary requests continue at normal latency.

Architectural principles

  • Nothing an attacker can trigger may run long on the event loop. CPU-heavy work moves to a pool that actually runs in parallel with the loop — and that depends on whether the code releases the GIL.
  • Every client-controlled multiplier has a bound: list lengths, connection lifetimes, body sizes, concurrent operations.
  • Authenticate exactly what was sent: signatures over raw bytes, compared in constant time, with replay windows.
  • Cache security material carefully: single-flight refreshes, rate-limited misses, last-good fallback.
  • Measure under attack conditions, not just normal load, because averages hide every one of these failures.
Where asyncio-specific attacks land 5 stacked layers. Where asyncio-specific attacks land connections slow clients exhaust file descriptors requests huge bodies and lists multiply work CPU on the loop hashing, backtracking regexes stall everyone authentication signatures over the wrong bytes, key-fetch floods upstreams fan-out sized by the client Each layer has one guide in this section.

Execution model: why one thread changes the threat model

In a thread-per-request server, a slow request occupies one thread; the others keep serving. In asyncio, request handlers are coroutines sharing one thread, and they cooperate by yielding at await. A handler that does 170 ms of CPU work without yielding — one bcrypt check — delays every other request on the worker by 170 ms; thirty-two of them in a burst delayed everything by 5.4 seconds in testing. The cheapest attack on an asyncio service is therefore not volume but CPU: find an endpoint that does something expensive per request and call it a little more than the service can absorb.

Offloading is the standard answer, and it only works when the offloaded code can run while the loop runs. C extensions that release the GIL — bcrypt, argon2-cffi, most compression and hashing libraries — can be offloaded to threads. Code that holds the GIL — Python's re matching, pure-Python parsers — cannot; in a thread it still competes with the loop for the interpreter, which is why the regex measurement stalled the loop as badly from a thread as from the loop itself. Those need a process pool, or better, a rewrite that removes the cost. On the I/O side, asyncio's efficiency with idle connections means a slow-client attack does not exhaust memory or threads, but it does exhaust file descriptors, and nothing in the event loop notices until accept fails.

The measurements behind this section A grid of 6 rows by 3 columns. The measurements behind this section attack or mistake measured guide password hashing on the loop 5.4 s stall for 32 logins hashing backtracking regex in a thread 795 ms loop stall (re holds the GIL) regex DoS slow clients vs fd limit 1,024 0 of 20 real requests served slow clients HMAC over re-serialized JSON valid signature rejected webhooks cold JWKS cache, 200 requests 200 fetches; single-flight: 1 JWKS 50,000 IDs in one request 50,000 concurrent calls; capped: rejected in 12.6 ms fan-out Python 3.14; each figure comes from the guide for that problem.

Pattern catalogue

CPU-heavy security work off the loop

Password hashing in a dedicated, bounded thread pool — bcrypt and argon2-cffi release the GIL — with a semaphore that rejects excess attempts. bcrypt at cost 12 cost 171 ms per check; in threads it reached 79.7 logins per second with loop lag p99 of 21 ms, and an eight-process pool kept lag at 1.1 ms. Hashing a dummy password for unknown usernames keeps response times from revealing which accounts exist. See hashing passwords without blocking the event loop.

Patterns that cannot backtrack

Atomic groups and possessive quantifiers (Python 3.11+), non-overlapping structure, input length caps, and a process pool with a timeout for patterns you cannot change. The possessive rewrite of ^(a+)+$ matched every test input in 0.8–2 µs, against 1.7 s for the original on 27 characters; a process pool kept the loop responsive but still spent the 772 ms of CPU in a worker, so it is a containment measure, not a fix. See preventing regex denial of service in async handlers.

Connections that cannot be held forever

A reverse proxy with header and body timeouts, a measured server timeout for incomplete requests, and a high, monitored descriptor limit. In testing, uvicorn's keep-alive timeout and --limit-concurrency did not close connections still sending their first headers, while Hypercorn's keep-alive timeout did and kept serving under the same attack; behaviour like this varies by server and version, so the attack itself belongs in the test suite. See defending async servers against slow clients.

Signatures over raw bytes

await request.body(), the provider's exact HMAC format, hmac.compare_digest, a timestamp window and event-ID deduplication. HMAC cost 1.4 µs per kilobyte, so verification belongs on the loop; the risks are verifying the wrong bytes and comparing with ==, which took measurably different times depending on where the strings differed. See verifying webhook signatures in async handlers.

A key cache that resists abuse

Single-flight JWKS refreshes, a minimum interval between refreshes triggered by unknown key IDs, background refresh with jitter, and last-good keys on failure. Verifying a token itself was cheap — 36 µs for RS256 — so the risk sits in the key cache: on a cold cache 200 requests made 200 fetches without single-flight and 1 with it, and 1,000 forged key IDs made 10 fetches with single-flight alone and 1 with a 30-second minimum interval. See validating JWTs with cached JWKS in asyncio.

Bounded fan-out

Length caps on request lists, shared per-upstream semaphores, bulk calls instead of per-item calls. The cap rejected a 50,000-ID request in 12.6 ms with no upstream calls, and the semaphore held a legitimate 100-ID request to 20 concurrent calls. See bounding user-controlled fan-out.

Testing under attack conditions

Each problem in this section is invisible under ordinary load tests, which send well-formed, average-sized requests at a steady rate. Testing for them means sending the inputs an attacker would. A short adversarial suite, run against staging, covers most of it: a burst of logins well above normal concurrency while measuring loop lag on the same worker; every regex-validated field filled with long near-matching strings; a few thousand connections that open and never finish their headers, with an ordinary request issued every second; webhook deliveries with modified bodies, old timestamps and repeated event IDs; tokens with random key IDs; and batch requests at, and just above, every documented limit.

What to measure is the same in every case: whether ordinary requests keep succeeding at normal latency while the adversarial traffic runs, and how quickly the hostile requests are rejected. A service that survives the suite rejects bad input in milliseconds, keeps loop lag flat, and keeps serving everyone else. The tools are the ones used throughout this section — a lag probe, a descriptor count, an upstream-call counter — and the load testing guidance applies, with one difference: here the interesting numbers are the maxima, not the percentiles.

Which defence does this input need? A decision on What does the attacker control with 4 outcomes. Which defence does this input need? What does the attacker control? how much CPU a request costs pool that releases the GIL, or rewrite 5.4 s -> 21 ms lag how many items or bytes caps + shared semaphores rejected in ms how long a connection stays timeouts per phase + fd budget 0/20 served otherwise what it claims to be raw-byte HMAC, guarded key cache 200 -> 1 fetch Bound the cost of every request before it starts.

Resource boundaries

  • Hashing: a dedicated pool of a few workers and a semaphore of a few dozen pending attempts per instance, plus per-account and per-IP rate limits.
  • Regexes: field lengths capped before matching; any validation over 10 ms treated as a bug.
  • Connections: a header timeout of about 10 seconds at the proxy; a descriptor limit far above peak connections, alerted at 70%.
  • Bodies: 1 MB or less on unauthenticated endpoints such as webhooks, enforced while streaming.
  • Key refreshes: at most one per 30 seconds triggered by unknown key IDs; background refresh every few minutes with jitter.
  • Fan-out: a documented maximum per request, and a shared semaphore per upstream sized from its capacity.
  • Error detail: generic messages to clients ("invalid token", "invalid credentials") with the specific reason in logs, so rejections do not teach an attacker which check failed.

Integrated production example

A FastAPI application combining the defences: bounded password hashing with the same cost for unknown users, a username validator with no nested quantifiers and a length cap, a webhook endpoint that verifies a timestamped HMAC over the raw body with a size limit and deduplication, and a batch endpoint with a length cap and a shared upstream semaphore. Tested with FastAPI's test client: a valid login took 186 ms and an unknown username 173 ms; a hostile 31-character username and a 101-ID batch were rejected with 422; a 100-ID batch completed in 55 ms; the webhook accepted a signed event, reported the repeat as a duplicate and rejected a tampered body with 401; and 16 concurrent logins all succeeded in 0.40 s:

import asyncio
import hashlib
import hmac
import json
import re
import time
from concurrent.futures import ThreadPoolExecutor

import bcrypt
from fastapi import FastAPI, HTTPException, Request
from pydantic import BaseModel, Field, field_validator

app = FastAPI()
HASH_POOL = ThreadPoolExecutor(max_workers=8, thread_name_prefix="pwhash")
HASH_SLOTS = asyncio.Semaphore(32)
CATALOG_SLOTS = asyncio.Semaphore(20)
WEBHOOK_SECRET = b"whsec_example"
USERNAME = re.compile(r"^[a-z0-9](?:[a-z0-9_-]{0,30}[a-z0-9])?$")   # no nested quantifiers
DUMMY_HASH = bcrypt.hashpw(b"dummy", bcrypt.gensalt(rounds=12))


class Login(BaseModel):
    username: str = Field(max_length=32)
    password: str = Field(max_length=256)

    @field_validator("username")
    @classmethod
    def check_username(cls, v: str) -> str:
        if not USERNAME.fullmatch(v):
            raise ValueError("invalid username")
        return v


class Batch(BaseModel):
    ids: list[int] = Field(min_length=1, max_length=100)


@app.post("/login")
async def login(body: Login):
    if HASH_SLOTS.locked():
        raise HTTPException(429, "too many login attempts in progress")
    stored = await users.password_hash(body.username) or DUMMY_HASH   # same cost for unknown users
    async with HASH_SLOTS:
        loop = asyncio.get_running_loop()
        ok = await loop.run_in_executor(HASH_POOL, bcrypt.checkpw, body.password.encode(), stored)
    if not ok or stored is DUMMY_HASH:
        raise HTTPException(401, "invalid credentials")
    return {"user": body.username}


@app.post("/webhooks/payments")
async def payments(request: Request):
    if int(request.headers.get("content-length", 0)) > 1 << 20:
        raise HTTPException(413, "payload too large")
    body = await request.body()
    ts = int(request.headers.get("x-timestamp", "0"))
    if abs(time.time() - ts) > 300:
        raise HTTPException(400, "stale timestamp")
    expected = hmac.new(WEBHOOK_SECRET, f"{ts}.".encode() + body, hashlib.sha256).hexdigest()
    if not hmac.compare_digest(expected, request.headers.get("x-signature", "")):
        raise HTTPException(401, "bad signature")
    event = json.loads(body)
    if not await processed.add_if_absent(event["id"], ttl=600):
        return {"duplicate": True}
    return {"accepted": event["id"]}


async def fetch_item(item_id: int) -> dict:
    async with CATALOG_SLOTS:
        return await catalog.get(item_id)


@app.post("/items/batch")
async def batch(body: Batch):
    async with asyncio.TaskGroup() as tg:
        tasks = [tg.create_task(fetch_item(i)) for i in body.ids]
    return [t.result() for t in tasks]

The tested version used an in-memory user table, a set for deduplication and a 10 ms sleep for the catalog; production replaces them with the user store, Redis SET NX and the real client, without changing the structure. JWT verification would sit in a shared dependency as shown in the JWKS guide, and the slow-client defence lives in front of the application, in the proxy.

Diagnostic hook callout

The metrics that reveal these attacks are maxima and counts, not averages:

  • Maximum event-loop lag per minute — a single multi-second stall points at CPU work on the loop (a hash, a regex) and the request that triggered it.
  • Open file descriptors against the limit — rising without a matching rise in completed requests means slow clients or leaks.
  • Maximum upstream calls per request — any request far above the documented cap means an unbounded fan-out path.
  • Rejections by reason — bad signatures, stale timestamps, unknown key IDs, oversized inputs; a sudden rise in one is either an attack or a regression.

Alert on loop lag above 250 ms, descriptors above 70% of the limit, any request exceeding its fan-out cap, and JWKS fetches above a handful per hour per instance.

Failure modes

Failure mode Root cause Detection Fix
Whole worker stalls during logins Hashing on the event loop Max loop lag spikes with logins Dedicated, bounded hash pool
Stall persists after moving to a thread Code holds the GIL (re) Loop lag unchanged from a thread Rewrite the pattern; process pool
Real requests fail, CPU idle Descriptors held by slow clients fds near limit, no completed requests Proxy header timeouts; higher limit
Valid webhooks rejected HMAC over parsed/re-serialized data Signature failures after a change Verify raw bytes
Identity provider flooded Uncoordinated or unbounded JWKS refreshes Fetches per instance Single-flight + minimum interval
One request floods an upstream Fan-out sized by input Upstream calls per request Caps, semaphores, bulk calls

Frequently Asked Questions

What security issues are specific to asyncio services?

Anything that holds the single event-loop thread — password hashing, backtracking regexes — stalls every request on the worker, and cheap concurrency lets clients multiply work through fan-out or held connections. In testing, 32 concurrent bcrypt logins on the loop stalled it for 5.4 s.

Is moving CPU work to a thread always enough in asyncio?

No. It works for C code that releases the GIL, such as bcrypt; Python's re holds the GIL, and a slow regex in a thread still stalled the loop for 795 ms in testing. Use a process pool or rewrite the work.

Are asyncio servers immune to slowloris?

They handle idle connections cheaply, but each holds a file descriptor: 1,100 slow clients exhausted a 1,024 limit in uvicorn and every real request failed. Put a proxy with header timeouts in front.

How do I verify webhooks in an async framework?

Compute the HMAC over the raw request bytes, compare with hmac.compare_digest, reject stale timestamps and deduplicate event IDs; verifying re-serialized JSON rejected a valid signature in testing.

How do I stop one request from overwhelming an upstream?

Cap list lengths at the request boundary and share a semaphore per upstream across requests: a 50,000-ID request was rejected in 12.6 ms instead of starting 50,000 concurrent calls.