Skip to content

Limiting Response Size in Async HTTP Clients

Neither httpx nor aiohttp limits how large a response body can be, so a client that calls .content or .read() on an untrusted URL lets the server decide how much memory the process uses. Compressed responses make it worse, because the limit that matters is the decompressed size. Measured on Python 3.14 with httpx 0.28.1 and aiohttp 3.14.3 against a local server: reading a 300 MB response the default way peaked at 647–756 MiB of resident memory. A 1.0 MB gzip response that decompressed to 1 GiB peaked at 2,108–2,125 MiB in both clients. Streaming with a 10 MiB limit on decoded bytes stopped every case: a declared Content-Length of 300 MB was refused before reading the body, at 47–57 MiB peak, and the gzip bomb was cut off at 58 MiB with aiohttp — but at 185 MiB with httpx, because one compressed network chunk expanded to tens of megabytes before the loop could check it. This guide builds the limit for both clients.

Prerequisites

1. Measure what an unbounded read costs

The convenient calls read the whole body before returning:

async with httpx.AsyncClient() as c:
    body = (await c.get(url)).content           # whole body in memory

async with aiohttp.ClientSession() as s:
    async with s.get(url) as r:
        body = await r.read()                   # same

Measured against three endpoints — 300 MB with chunked encoding, 300 MB with a Content-Length header, and a 1,043,656-byte gzip body of 1 GiB of zeros — each in a fresh process: httpx peaked at 694 MiB, 756 MiB and 2,125 MiB; aiohttp at 647 MiB, 647 MiB and 2,108 MiB. Peak memory was about twice the body size, as buffers were copied while assembling it. A service that fetches user-supplied URLs — webhooks, link previews, imports — can be taken down by one such response, and by a few concurrent ones even if each alone fits.

Verify: for each client call on an untrusted URL, the code bounds the body size before or while reading it.

Peak memory reading one response A grid of 3 rows by 5 columns. Peak memory reading one response response httpx default httpx, 10 MiB limit aiohttp default aiohttp, 10 MiB limit 300 MB, chunked 694 MiB 67 MiB 647 MiB 58 MiB 300 MB, Content-Length 756 MiB 57 MiB, refused early 647 MiB 47 MiB, refused early 1.0 MB gzip -> 1 GiB 2,125 MiB 185 MiB 2,108 MiB 58 MiB Peak RSS of a fresh process; httpx 0.28.1, aiohttp 3.14.3.

2. Refuse early on Content-Length

When the server declares the size, check it before reading anything:

class ResponseTooLarge(Exception):
    pass

LIMIT = 10 * 1024 * 1024

async with client.stream("GET", url) as r:
    declared = r.headers.get("content-length")
    if declared is not None and int(declared) > LIMIT:
        raise ResponseTooLarge(f"declared {declared} bytes")

Measured: both clients refused the 300 MB Content-Length response with a peak of 47–57 MiB — the interpreter and libraries — and in under 0.1 s. This check is cheap and catches honest servers sending large files, but it is not a limit by itself: a server can omit the header and use chunked encoding, and with Content-Encoding: gzip the header gives the compressed size, which says nothing about the decoded size.

Verify: a response with a large declared length is refused before its body is read, in milliseconds.

3. Count decoded bytes while streaming

The real limit counts bytes after decompression, as they arrive:

async def fetch_limited_httpx(client: httpx.AsyncClient, url: str, limit=LIMIT) -> bytes:
    async with client.stream("GET", url) as r:
        r.raise_for_status()
        declared = r.headers.get("content-length")
        if declared is not None and int(declared) > limit:
            raise ResponseTooLarge(f"declared {declared}")
        body = bytearray()
        async for chunk in r.aiter_bytes():           # decoded, decompressed bytes
            body += chunk
            if len(body) > limit:
                raise ResponseTooLarge(f"over {limit} bytes")
        return bytes(body)

async def fetch_limited_aiohttp(session: aiohttp.ClientSession, url: str, limit=LIMIT) -> bytes:
    async with session.get(url) as r:
        r.raise_for_status()
        if r.content_length is not None and r.content_length > limit:
            raise ResponseTooLarge(f"declared {r.content_length}")
        body = bytearray()
        async for chunk in r.content.iter_chunked(64 * 1024):
            body += chunk
            if len(body) > limit:
                raise ResponseTooLarge(f"over {limit} bytes")
        return bytes(body)

Measured: the 300 MB chunked response was stopped at 67 MiB peak with httpx and 58 MiB with aiohttp. Leaving the async with block closes the response, so the rest of the body is never read: the abort took 0.01–0.10 s for a 300 MB body. Use aiter_bytes, not aiter_raw: raw bytes are the compressed ones, and counting them lets a bomb through.

Verify: a chunked response without Content-Length and larger than the limit raises ResponseTooLarge, and peak memory stays near the limit.

Bounding a response body A flow of 5 stages. Bounding a response body Open as a stream headers only so far Content-Length > limit? refuse, no body read Iterate decoded chunks count after decompression Count > limit? raise, leave the block Response closed rest never read The decoded count catches chunked bodies and compression bombs.

4. Bound the size of one decoded chunk

The gzip bomb exposed a difference between the clients. With aiohttp, iter_chunked(64 * 1024) yields at most 64 KiB at a time, so the loop saw the limit crossed almost immediately: 58 MiB peak, 0.03 s. With httpx, aiter_bytes() yields whatever one network read decompresses to, and a few kilobytes of gzipped zeros decompress to many megabytes: the peak was 185 MiB and the abort took 0.31 s. Passing a chunk size re-slices the output, but the decoder has already produced the whole expansion:

async for chunk in r.aiter_bytes(chunk_size=64 * 1024):   # smaller pieces, same decoder output
    ...

For a hard bound on a compression bomb with httpx, decompress yourself with a size cap — read aiter_raw() and feed it to zlib.decompressobj with max_length, which never produces more than the requested amount per call — or refuse compressed responses from untrusted sources by sending Accept-Encoding: identity. The 185 MiB peak is still far better than 2.1 GiB; whether it is acceptable depends on how many such fetches run concurrently.

Verify: a gzip bomb test against the production client code shows peak memory within budget, at the expected concurrency.

Peak memory, 1.0 MB gzip body decoding to 1 GiB 4 horizontal bars comparing httpx, .content with the others. Peak memory, 1.0 MB gzip body decoding to 1 GiB httpx, .content 2,125 MiB aiohttp, .read() 2,108 MiB httpx, aiter_bytes + limit 185 MiB aiohttp, iter_chunked(64 KiB) + limit 58 MiB One network chunk can expand a thousandfold before the check runs.

5. Combine size limits with time limits

A size limit stops large bodies; a server that sends a small body very slowly needs a time limit too. Wrap the whole fetch, not just the connect or each read:

async def fetch_untrusted(client, url):
    async with asyncio.timeout(15):                 # total time, including a slow drip
        return await fetch_limited_httpx(client, url, limit=10 * 1024 * 1024)

A per-read timeout resets on every byte, so a server sending one byte every few seconds never trips it; a total deadline does. For concurrent fetches of untrusted URLs, the memory budget is the limit multiplied by the concurrency, so bound concurrency with a semaphore as well. Timeouts for each phase are covered in setting connect, read and total timeouts in async HTTP clients.

Verify: every fetch of an untrusted URL has a decoded-size limit, a total deadline, and a concurrency bound, and their product fits the memory budget.

Verification

Response size is bounded when:

  • No .content, .json() or .read() is called on an untrusted response without a prior limit.
  • Content-Length above the limit is refused before the body is read.
  • Decoded bytes are counted while streaming, so chunked bodies and compression bombs are stopped.
  • A total deadline and a concurrency bound complete the budget.

Diagnostic Hook: when a fetcher process's memory jumps by gigabytes on one request, check that response's Content-Encoding and compressed size. A 1.0 MB gzip body took 2.1 GiB to read here; the access log's byte count will look harmless.

Pitfalls & edge cases

  • Trusting Content-Length. Chunked responses omit it; gzip makes it the compressed size.
  • Counting raw bytes. aiter_raw counts compressed bytes and lets a bomb through.
  • Large decoded chunks. httpx peaked at 185 MiB on the bomb with a 10 MiB limit.
  • Size limits without time limits. A slow drip of small chunks never trips a size check.

Frequently Asked Questions

Does httpx or aiohttp limit response size?

No. Both read any body into memory: 300 MB took up to 756 MiB, and a 1 MB gzip bomb took 2.1 GiB. Stream the body and count bytes yourself.

How do I limit response size in httpx?

Use client.stream, refuse a large Content-Length, then iterate aiter_bytes and raise once the decoded total passes the limit. Leaving the block closes the response.

Am I protected from gzip bombs by a byte limit?

Only if you count decoded bytes. With aiohttp's 64 KiB chunks the bomb stopped at 58 MiB; httpx's decoder produced larger pieces, peaking at 185 MiB.

Is Content-Length enough to reject large responses?

No. It is absent with chunked encoding and is the compressed size with gzip. Use it to refuse early, then enforce the limit while streaming.