Skip to content

Limiting Request Body Size in ASGI Apps

Uvicorn has no request body size limit, and neither has Starlette or FastAPI: a handler that calls await request.body() or await request.json() will read whatever the client sends into memory. Tested on Uvicorn 0.54 and Starlette 1.7, a single 200 MB POST raised the worker's resident memory from 37 MB to 237 MB, and it was accepted with a 200. Ten such requests at once would take 2 GB. With a 10 MB limit enforced by a pure ASGI middleware, the same upload got a 413 in 146 ms when its Content-Length was declared, and in 15 ms when it was sent chunked with no length — the middleware counted bytes as they arrived and stopped at the limit. This guide builds that middleware, sets limits per route, and covers the proxy layer that should usually enforce it first.

Prerequisites

1. Reject declared oversize bodies before reading

Most clients send Content-Length. If it exceeds the limit, reject the request without reading a byte of the body:

class BodyLimit:
    def __init__(self, app, max_bytes: int) -> None:
        self.app, self.max = app, max_bytes

    async def __call__(self, scope, receive, send) -> None:
        if scope["type"] != "http":
            return await self.app(scope, receive, send)
        declared = dict(scope["headers"]).get(b"content-length")
        if declared is not None and int(declared) > self.max:
            return await self.reject(send)
        await self.app(scope, receive, send)       # step 2 adds the streaming check

    async def reject(self, send) -> None:
        await send({"type": "http.response.start", "status": 413,
                    "headers": [(b"content-length", b"0"), (b"connection", b"close")]})
        await send({"type": "http.response.body", "body": b""})

Connection: close tells the client the server will not read the rest of the body, so the connection cannot be reused for another request. A malformed Content-Length (not an integer) raises ValueError here; the HTTP parser in the server usually rejects those first, but catch it and answer 400 if your server passes them through.

Verify: a request with Content-Length: 209715200 gets a 413 and the handler never runs.

Worker memory after one 200 MB upload 3 horizontal bars comparing idle worker with the others. Worker memory after one 200 MB upload idle worker 37 MB 200 MB upload, no limit 237 MB, accepted 200 MB upload, 10 MB limit 37 MB, 413 Uvicorn 0.54, Starlette 1.7; handler calls await request.body(); RSS read from /proc. Without a limit, the body lands in memory in full — per concurrent request.

2. Count bytes for chunked bodies

A client can omit Content-Length and send the body with chunked transfer encoding, so the declared-size check alone is not enough. Wrap receive, count the bytes of each http.request message, and stop when the count passes the limit:

class TooLarge(Exception):
    pass


async def __call__(self, scope, receive, send) -> None:
    ...                                              # declared check as above
    seen = 0
    started = False

    async def limited_receive():
        nonlocal seen
        message = await receive()
        if message["type"] == "http.request":
            seen += len(message.get("body", b""))
            if seen > self.max:
                raise TooLarge                       # unwinds out of request.body()
        return message

    async def tracking_send(message):
        nonlocal started
        if message["type"] == "http.response.start":
            started = True
        await send(message)

    try:
        await self.app(scope, limited_receive, tracking_send)
    except TooLarge:
        if not started:
            await self.reject(send)

Measured: a 200 MB chunked upload was rejected with 413 after 15 ms, having read just over 10 MB. Raising from receive unwinds the handler wherever it is reading the body; tracking whether the response has started avoids sending a second http.response.start, which the server would reject. Memory stays bounded at the limit, not at the upload size.

Verify: stream a chunked body larger than the limit; the response is 413 and the worker's memory does not grow beyond the limit.

Two checks, one limit A flow of 4 stages. Two checks, one limit declared length? > limit: 413 now wrap receive count body bytes count > limit raise TooLarge response not started send 413, close Measured: 146 ms to reject a declared 200 MB body, 15 ms for a chunked one.

3. Set limits per route

A single global limit is either too small for the upload endpoint or too large for everything else. JSON APIs rarely need more than a megabyte; file uploads may need hundreds. Choose the limit from the path:

class RouteBodyLimit:
    def __init__(self, app, default: int, overrides: dict[str, int]) -> None:
        self.default = BodyLimit(app, default)
        self.overrides = [(prefix, BodyLimit(app, n)) for prefix, n in overrides.items()]

    async def __call__(self, scope, receive, send) -> None:
        limiter = self.default
        if scope["type"] == "http":
            limiter = next((lim for prefix, lim in self.overrides
                            if scope["path"].startswith(prefix)), self.default)
        await limiter(scope, receive, send)


app = Starlette(routes=routes, middleware=[
    Middleware(RouteBodyLimit, default=1 << 20,                 # 1 MiB for the API
               overrides={"/uploads/": 512 << 20}),             # 512 MiB for uploads
])

For the large-upload routes, do not buffer at all: read request.stream() and write chunks to disk or object storage as they arrive, as in uploading large files to S3 with async multipart. The limit then bounds the disk used, and memory stays at one chunk.

Verify: a 2 MB JSON body to an API route gets 413; a 100 MB file to the upload route succeeds with flat memory.

4. Enforce it at the proxy as well

The cheapest place to reject an oversized body is before it reaches Python. nginx's client_max_body_size (default 1 MB) rejects with 413 at the edge; most load balancers and API gateways have an equivalent. Keep both layers:

server {
    client_max_body_size 1m;           # API default
    location /uploads/ {
        client_max_body_size 512m;
        proxy_request_buffering off;   # stream to the app instead of spooling to disk first
        proxy_pass http://app;
    }
}

The proxy protects the app from bulk abuse; the app-level middleware protects it when it is reached some other way — another proxy configuration, direct access within the cluster, a test environment — and encodes the per-route intent next to the code. Keep the two limits consistent, so the app's 413 is a fallback rather than a surprise.

Verify: send an oversized request through the proxy and directly to the app; both answer 413.

5. Bound multipart and decompressed bodies too

Two cases slip past a raw byte limit. Multipart form parsing in Starlette spools uploaded files to temporary files, with its own per-part limits (max_part_size, max_files, max_fields on request.form()); set them deliberately. And a compressed body can be small on the wire and enormous when decompressed — limit the decompressed size wherever you decompress:

import zlib


def bounded_gunzip(data: bytes, max_out: int) -> bytes:
    d = zlib.decompressobj(16 + zlib.MAX_WBITS)
    out = d.decompress(data, max_out)                 # stops at max_out bytes
    if d.unconsumed_tail:
        raise TooLarge("decompressed body exceeds limit")
    return out


async def ingest(request):
    form = await request.form(max_files=5, max_fields=50, max_part_size=1 << 20)

The max_length argument of decompress stops output at the cap and leaves the rest in unconsumed_tail, so a gzip bomb is caught at the limit instead of exhausting memory: measured, a 1.04 MB gzip body that inflates to 1 GiB stopped at a 10 MiB cap.

Verify: a gzip body that inflates past the limit is rejected, and the worker's memory stays flat.

Where body limits belong 4 stacked layers. Where body limits belong proxy client_max_body_size, cheapest rejection ASGI middleware per-route limit, declared and chunked form parsing max_files, max_part_size decompression bounded output size Each layer catches a case the one above it cannot see.

Verification

Body limits work when:

  • Declared oversize bodies are rejected with 413 before any body is read.
  • Chunked bodies are rejected at the limit, with memory bounded by the limit.
  • Limits are per route, small for APIs and streamed-to-disk for uploads.
  • Proxy and app limits agree, and multipart and decompression are bounded too.

Diagnostic Hook: count 413 responses per route and layer (proxy versus app), and graph worker memory against request body sizes. 413s that come only from the app mean the proxy limit is missing or larger; memory spikes that line up with large requests on routes you thought were limited mean a code path reads the body before the middleware sees it.

Pitfalls & edge cases

  • Trusting Content-Length alone. Chunked uploads omit it; count streamed bytes.
  • Buffering large uploads. Stream them; the limit then bounds disk, not memory.
  • Sending 413 after the response started. Track it and skip the second start.
  • Unbounded decompression. Small compressed bodies can inflate past any limit.

Frequently Asked Questions

Does Uvicorn limit request body size?

No. In testing, Uvicorn 0.54 with Starlette accepted a 200 MB body into memory, raising worker memory from 37 MB to 237 MB. Enforce a limit at the proxy and in ASGI middleware.

How do I limit request body size in FastAPI?

Add pure ASGI middleware that rejects a too-large Content-Length with 413 and wraps receive to count streamed bytes, raising once the count passes the limit.

How do I handle chunked uploads without Content-Length?

Count the bytes of each http.request message as it arrives and stop at the limit. A chunked 200 MB upload was rejected after 15 ms with a 10 MB limit.

Should the body size limit be in nginx or the app?

Both. The proxy rejects cheaply at the edge; the app enforces per-route limits and protects paths that bypass the proxy.