Skip to content

Enforcing Request Timeouts in ASGI Servers

Uvicorn bounds idle keep-alive connections and graceful shutdown, but not how long a request may run: a handler that takes five minutes is allowed five minutes. Most services want a ceiling, so a stuck dependency turns into a fast, explicit error instead of piled-up requests. Tested with Uvicorn 0.54 and Starlette 1.7: without a timeout, a handler sleeping 5 s returned 200 after 5.06 s. Behind a pure ASGI middleware with a 1-second asyncio.timeout, the same request got 504 after 1.06 s, and the handler was cancelled at 1.00 s — the work stopped, not just the wait. A streaming response that had already started when the timeout fired could not be turned into a 504: the connection was dropped and the client saw RemoteProtocolError after 1.02 s, with a truncated body. This guide adds request timeouts that stop work, return clear errors, and treat streaming responses differently.

Prerequisites

1. Wrap the app in a timeout middleware

A pure ASGI middleware runs the app inside asyncio.timeout and, if the deadline passes before the response has started, sends a 504 itself:

class RequestTimeout:
    def __init__(self, app, seconds: float) -> None:
        self.app, self.seconds = app, seconds

    async def __call__(self, scope, receive, send) -> None:
        if scope["type"] != "http":
            return await self.app(scope, receive, send)
        started = False

        async def tracking_send(message) -> None:
            nonlocal started
            if message["type"] == "http.response.start":
                started = True
            await send(message)

        try:
            async with asyncio.timeout(self.seconds):
                await self.app(scope, receive, tracking_send)
        except TimeoutError:
            if started:
                raise                        # too late for a status code: drop the connection
            await send({"type": "http.response.start", "status": 504,
                        "headers": [(b"content-type", b"application/json")]})
            await send({"type": "http.response.body", "body": b'{"error":"request timed out"}'})

Measured: 504 after 1.06 s for a handler that would have taken 5 s, and the handler received CancelledError at 1.00 s — so its database query or HTTP call was cancelled too. A pure ASGI middleware does this cheaply; Starlette's BaseHTTPMiddleware adds overhead and runs the app in a separate task, which complicates cancellation. Put the timeout middleware outermost so it covers everything, including other middleware.

Verify: a test endpoint that sleeps past the limit returns 504 at the limit, and logs from the handler show it was cancelled.

The same requests with and without a 1 s timeout middleware A grid of 2 rows by 3 columns. The same requests with and without a 1 s timeout middleware endpoint no timeout 1 s timeout middleware /slow (5 s handler) 200 after 5.06 s 504 after 1.06 s; handler cancelled at 1.00 s /stream (10 chunks, 3 s) 200, complete, 3.02 s RemoteProtocolError at 1.02 s, truncated Uvicorn 0.54, Starlette 1.7; a started response cannot become a 504.

2. Treat streaming responses separately

Once http.response.start has been sent, the status is fixed. A timeout after that point can only abort the connection, which the client sees as a protocol error — tested: RemoteProtocolError after 1.02 s with a partial body. Streaming endpoints need a different rule:

class RequestTimeout:
    def __init__(self, app, seconds: float, exempt_prefixes: tuple[str, ...] = ()) -> None:
        self.app, self.seconds, self.exempt = app, seconds, exempt_prefixes

    async def __call__(self, scope, receive, send) -> None:
        if scope["type"] != "http" or scope["path"].startswith(self.exempt):
            return await self.app(scope, receive, send)       # streams manage their own limits
        ...


app = RequestTimeout(app, seconds=10, exempt_prefixes=("/exports/", "/events/"))

Time to first byte is still worth bounding for streams — a stream that never starts is as stuck as any other request — and the gap between chunks is the right ongoing limit, as in timing out each item of an async iterator. A whole-duration limit on a large download or a server-sent-events feed only guarantees that legitimate long responses are cut.

Verify: a long, healthy export completes, while an export whose first byte never comes fails at the first-byte limit.

3. Set different budgets per route

One limit for everything is either too long for fast endpoints or too short for slow ones. Configure budgets by route, and let handlers read the remaining time:

import contextvars

deadline: contextvars.ContextVar[float | None] = contextvars.ContextVar("deadline", default=None)

BUDGETS = {"/search": 2.0, "/reports": 30.0}
DEFAULT_BUDGET = 5.0


class RouteTimeout:
    def __init__(self, app) -> None:
        self.app = app

    async def __call__(self, scope, receive, send) -> None:
        if scope["type"] != "http":
            return await self.app(scope, receive, send)
        budget = next((s for p, s in BUDGETS.items() if scope["path"].startswith(p)), DEFAULT_BUDGET)
        token = deadline.set(asyncio.get_running_loop().time() + budget)
        try:
            await RequestTimeout(self.app, budget)(scope, receive, send)   # per-request budget
        finally:
            deadline.reset(token)


def remaining() -> float | None:
    d = deadline.get()
    return None if d is None else max(0.0, d - asyncio.get_running_loop().time())

Exposing the deadline in a context variable lets handlers pass the remaining budget to downstream calls (timeout=remaining()), so a request near its limit does not start a ten-second call it cannot wait for. Choose budgets from measured latency — for example a multiple of p99 — as in deriving timeouts from latency percentiles.

Verify: each route's budget appears in configuration with the latency data it came from, and downstream calls use remaining().

A request under a timeout middleware A flow of 5 stages. A request under a timeout middleware pick budget per route, deadline contextvar asyncio.timeout(app) track response start deadline passes handler cancelled not started -> 504 started -> drop downstream calls use remaining() Tested: 504 at 1.06 s, handler cancelled at 1.00 s.

4. Make cancelled handlers clean up correctly

A request timeout is a cancellation inside the handler, so everything about cancellation safety applies: transactions must roll back, side effects must be idempotent or shielded, and cleanup must not swallow the CancelledError:

async def create_order(request):
    payload = await request.json()
    async with pool.acquire() as conn, conn.transaction():      # rolled back if cancelled
        order_id = await conn.fetchval("insert into orders ... returning id", ...)
        await conn.execute("insert into outbox ...", ...)      # published after commit
    return JSONResponse({"id": order_id}, status_code=201)

If the deadline passes during the transaction, it rolls back and the client gets 504; if it passes after the commit but before the response is sent, the order exists and the client still gets 504 — so clients must be able to retry safely with an idempotency key, as covered in making database transactions cancellation-safe. The 504 tells the client "unknown outcome, retry safely", not "nothing happened".

Verify: a request that times out after committing can be retried with the same idempotency key without creating a duplicate.

5. Align with proxy and client timeouts

The service's request timeout is one of several on the path. They should nest, outermost longest:

# client timeout        > load balancer / proxy timeout  > app request timeout  > downstream call timeouts
CLIENT_TIMEOUT = 30          # what callers configure
PROXY_READ_TIMEOUT = 25      # e.g. nginx proxy_read_timeout
APP_REQUEST_TIMEOUT = 20     # this middleware
DOWNSTREAM_TIMEOUT = 5       # per call, using remaining()

If the proxy times out before the app, the client gets the proxy's generic 504 while the app keeps working on a request nobody will receive. If the app times out first, it can return a meaningful error and stop its work. The same nesting rule applies one level down: downstream calls get less than the request's remaining budget. Log 504s with the route and the step that was in progress (the CancelledError traceback from the handler shows it), and alert on their rate per route.

Verify: during a forced slowdown of a dependency, clients receive the app's 504 body, not the proxy's, and the app's logs name the slow step.

What timeout should this route get? A decision on What kind of route is it with 4 outcomes. What timeout should this route get? What kind of route is it? request-response middleware timeout -> 504 handler cancelled streaming / downloads exempt; first-byte + chunk gaps no mid-body cuts has side effects + idempotency keys 504 means unknown outcome behind a proxy app < proxy < client meaningful 504s Bound requests in the app, where the error can be meaningful and the work can stop.

Verification

Request timeouts are enforced well when:

  • A pure ASGI middleware bounds every request-response route and returns 504 before the response starts.
  • Streaming routes are exempt from total limits and bound first byte and chunk gaps instead.
  • Budgets are per route, exposed to handlers for downstream calls.
  • Handlers are cancellation-safe and timeouts nest within proxy and client limits.

Diagnostic Hook: count 504s per route and capture the handler's CancelledError traceback for each. The innermost frame shows where the request was waiting when the budget ran out — a specific query, a downstream service — which is the thing to fix or to give its own shorter timeout.

Pitfalls & edge cases

  • Relying on Uvicorn for request limits. Tested: a 5 s handler ran 5.06 s.
  • Total timeouts on streaming routes. Tested: the response was cut mid-body.
  • Proxy timeouts shorter than the app's. Clients get generic errors and the app keeps working.
  • Treating 504 as "nothing happened". The commit may have succeeded.

Frequently Asked Questions

Does Uvicorn have a request timeout?

No. It has keep-alive and graceful-shutdown timeouts, but a slow handler runs as long as it takes; in testing a 5-second handler returned after 5.06 s.

How do I add a request timeout to FastAPI or Starlette?

Add pure ASGI middleware that runs the app inside asyncio.timeout and sends a 504 if the response has not started. In testing it returned 504 at 1.06 s and cancelled the handler at 1.00 s.

What happens to a streaming response when a request timeout fires?

The status has already been sent, so the server can only drop the connection; the client saw RemoteProtocolError and a truncated body in testing. Exempt streaming routes and bound first byte and chunk gaps instead.

Does a request timeout stop the handler's work?

With asyncio.timeout around the app, yes: the handler is cancelled, and awaited database queries and HTTP calls are cancelled with it.