Enforcing Request Timeouts in ASGI Servers¶
Uvicorn bounds idle keep-alive connections and graceful shutdown, but not how long a request may run: a handler that takes five minutes is allowed five minutes. Most services want a ceiling, so a stuck dependency turns into a fast, explicit error instead of piled-up requests. Tested with Uvicorn 0.54 and Starlette 1.7: without a timeout, a handler sleeping 5 s returned 200 after 5.06 s. Behind a pure ASGI middleware with a 1-second asyncio.timeout, the same request got 504 after 1.06 s, and the handler was cancelled at 1.00 s — the work stopped, not just the wait. A streaming response that had already started when the timeout fired could not be turned into a 504: the connection was dropped and the client saw RemoteProtocolError after 1.02 s, with a truncated body. This guide adds request timeouts that stop work, return clear errors, and treat streaming responses differently.
Prerequisites¶
- Python 3.11+, Starlette or FastAPI on Uvicorn (tested with 1.7 and 0.54).
- Pure ASGI middleware, from writing pure ASGI middleware.
- Deadlines across services, from propagating deadlines across async service calls.
1. Wrap the app in a timeout middleware¶
A pure ASGI middleware runs the app inside asyncio.timeout and, if the deadline passes before the response has started, sends a 504 itself:
class RequestTimeout:
def __init__(self, app, seconds: float) -> None:
self.app, self.seconds = app, seconds
async def __call__(self, scope, receive, send) -> None:
if scope["type"] != "http":
return await self.app(scope, receive, send)
started = False
async def tracking_send(message) -> None:
nonlocal started
if message["type"] == "http.response.start":
started = True
await send(message)
try:
async with asyncio.timeout(self.seconds):
await self.app(scope, receive, tracking_send)
except TimeoutError:
if started:
raise # too late for a status code: drop the connection
await send({"type": "http.response.start", "status": 504,
"headers": [(b"content-type", b"application/json")]})
await send({"type": "http.response.body", "body": b'{"error":"request timed out"}'})
Measured: 504 after 1.06 s for a handler that would have taken 5 s, and the handler received CancelledError at 1.00 s — so its database query or HTTP call was cancelled too. A pure ASGI middleware does this cheaply; Starlette's BaseHTTPMiddleware adds overhead and runs the app in a separate task, which complicates cancellation. Put the timeout middleware outermost so it covers everything, including other middleware.
Verify: a test endpoint that sleeps past the limit returns 504 at the limit, and logs from the handler show it was cancelled.
2. Treat streaming responses separately¶
Once http.response.start has been sent, the status is fixed. A timeout after that point can only abort the connection, which the client sees as a protocol error — tested: RemoteProtocolError after 1.02 s with a partial body. Streaming endpoints need a different rule:
class RequestTimeout:
def __init__(self, app, seconds: float, exempt_prefixes: tuple[str, ...] = ()) -> None:
self.app, self.seconds, self.exempt = app, seconds, exempt_prefixes
async def __call__(self, scope, receive, send) -> None:
if scope["type"] != "http" or scope["path"].startswith(self.exempt):
return await self.app(scope, receive, send) # streams manage their own limits
...
app = RequestTimeout(app, seconds=10, exempt_prefixes=("/exports/", "/events/"))
Time to first byte is still worth bounding for streams — a stream that never starts is as stuck as any other request — and the gap between chunks is the right ongoing limit, as in timing out each item of an async iterator. A whole-duration limit on a large download or a server-sent-events feed only guarantees that legitimate long responses are cut.
Verify: a long, healthy export completes, while an export whose first byte never comes fails at the first-byte limit.
3. Set different budgets per route¶
One limit for everything is either too long for fast endpoints or too short for slow ones. Configure budgets by route, and let handlers read the remaining time:
import contextvars
deadline: contextvars.ContextVar[float | None] = contextvars.ContextVar("deadline", default=None)
BUDGETS = {"/search": 2.0, "/reports": 30.0}
DEFAULT_BUDGET = 5.0
class RouteTimeout:
def __init__(self, app) -> None:
self.app = app
async def __call__(self, scope, receive, send) -> None:
if scope["type"] != "http":
return await self.app(scope, receive, send)
budget = next((s for p, s in BUDGETS.items() if scope["path"].startswith(p)), DEFAULT_BUDGET)
token = deadline.set(asyncio.get_running_loop().time() + budget)
try:
await RequestTimeout(self.app, budget)(scope, receive, send) # per-request budget
finally:
deadline.reset(token)
def remaining() -> float | None:
d = deadline.get()
return None if d is None else max(0.0, d - asyncio.get_running_loop().time())
Exposing the deadline in a context variable lets handlers pass the remaining budget to downstream calls (timeout=remaining()), so a request near its limit does not start a ten-second call it cannot wait for. Choose budgets from measured latency — for example a multiple of p99 — as in deriving timeouts from latency percentiles.
Verify: each route's budget appears in configuration with the latency data it came from, and downstream calls use remaining().
4. Make cancelled handlers clean up correctly¶
A request timeout is a cancellation inside the handler, so everything about cancellation safety applies: transactions must roll back, side effects must be idempotent or shielded, and cleanup must not swallow the CancelledError:
async def create_order(request):
payload = await request.json()
async with pool.acquire() as conn, conn.transaction(): # rolled back if cancelled
order_id = await conn.fetchval("insert into orders ... returning id", ...)
await conn.execute("insert into outbox ...", ...) # published after commit
return JSONResponse({"id": order_id}, status_code=201)
If the deadline passes during the transaction, it rolls back and the client gets 504; if it passes after the commit but before the response is sent, the order exists and the client still gets 504 — so clients must be able to retry safely with an idempotency key, as covered in making database transactions cancellation-safe. The 504 tells the client "unknown outcome, retry safely", not "nothing happened".
Verify: a request that times out after committing can be retried with the same idempotency key without creating a duplicate.
5. Align with proxy and client timeouts¶
The service's request timeout is one of several on the path. They should nest, outermost longest:
# client timeout > load balancer / proxy timeout > app request timeout > downstream call timeouts
CLIENT_TIMEOUT = 30 # what callers configure
PROXY_READ_TIMEOUT = 25 # e.g. nginx proxy_read_timeout
APP_REQUEST_TIMEOUT = 20 # this middleware
DOWNSTREAM_TIMEOUT = 5 # per call, using remaining()
If the proxy times out before the app, the client gets the proxy's generic 504 while the app keeps working on a request nobody will receive. If the app times out first, it can return a meaningful error and stop its work. The same nesting rule applies one level down: downstream calls get less than the request's remaining budget. Log 504s with the route and the step that was in progress (the CancelledError traceback from the handler shows it), and alert on their rate per route.
Verify: during a forced slowdown of a dependency, clients receive the app's 504 body, not the proxy's, and the app's logs name the slow step.
Verification¶
Request timeouts are enforced well when:
- A pure ASGI middleware bounds every request-response route and returns 504 before the response starts.
- Streaming routes are exempt from total limits and bound first byte and chunk gaps instead.
- Budgets are per route, exposed to handlers for downstream calls.
- Handlers are cancellation-safe and timeouts nest within proxy and client limits.
Diagnostic Hook: count 504s per route and capture the handler's CancelledError traceback for each. The innermost frame shows where the request was waiting when the budget ran out — a specific query, a downstream service — which is the thing to fix or to give its own shorter timeout.
Pitfalls & edge cases¶
- Relying on Uvicorn for request limits. Tested: a 5 s handler ran 5.06 s.
- Total timeouts on streaming routes. Tested: the response was cut mid-body.
- Proxy timeouts shorter than the app's. Clients get generic errors and the app keeps working.
- Treating 504 as "nothing happened". The commit may have succeeded.
Frequently Asked Questions¶
Does Uvicorn have a request timeout?
No. It has keep-alive and graceful-shutdown timeouts, but a slow handler runs as long as it takes; in testing a 5-second handler returned after 5.06 s.
How do I add a request timeout to FastAPI or Starlette?
Add pure ASGI middleware that runs the app inside asyncio.timeout and sends a 504 if the response has not started. In testing it returned 504 at 1.06 s and cancelled the handler at 1.00 s.
What happens to a streaming response when a request timeout fires?
The status has already been sent, so the server can only drop the connection; the client saw RemoteProtocolError and a truncated body in testing. Exempt streaming routes and bound first byte and chunk gaps instead.
Does a request timeout stop the handler's work?
With asyncio.timeout around the app, yes: the handler is cancelled, and awaited database queries and HTTP calls are cancelled with it.
Related¶
- Timeouts & Deadlines — up to the topic overview.
- Implementing idle timeouts for connections — the connection-level counterpart.
- Resilience, Cancellation & Error Handling — the section overview.