Skip to content

Compressing ASGI Responses Without Blocking the Loop

Response compression is CPU work, and in an ASGI app it runs wherever the middleware puts it. Starlette 1.7's GZipMiddleware already moves large bodies to worker threads — but only bodies of 128 KiB or more, at compression level 9 by default. Measured on Python 3.14 with Starlette 1.7.0 and uvicorn 0.54 on a 24-core machine, 16 clients requesting JSON with Accept-Encoding: gzip, and a separate probe timing a trivial /ping route: for a 117 KiB response — just under the threshold — the default middleware compressed on the event loop, served 291 responses per second, and the probe's p99 rose to 105 ms. Lowering thread_minimum_size to 0 moved the work to threads: 3,247 responses per second and a probe p99 of 1.9 ms. For a 1.1 MB response, the default level 9 gave 377 responses per second at 192.5 KiB each, against 1,136 at level 6 for 198.0 KiB — and the single worker process used 15 cores doing it. Serving a precompressed body reached 15,952 per second on under one core. This guide measures each option and picks settings.

Prerequisites

1. Measure compression cost per body

Before configuring middleware, time the codecs on a representative response. For a 1.1 MB JSON array of 10,000 user records:

for name, fn in [("gzip 9", lambda d: gzip.compress(d, 9)),
                 ("gzip 6", lambda d: gzip.compress(d, 6)),
                 ("gzip 1", lambda d: gzip.compress(d, 1)),
                 ("brotli 4", lambda d: brotli.compress(d, quality=4)),
                 ("brotli 11", lambda d: brotli.compress(d, quality=11))]:
    t = time.perf_counter()
    out = fn(data)
    print(name, (time.perf_counter() - t) * 1000, len(out))

Measured: gzip level 9 took 35.4 ms and produced 192.5 KiB; level 6, 9.8 ms and 198.0 KiB; level 1, 4.3 ms and 226.7 KiB. Brotli quality 4 took 5.7 ms for 150.4 KiB; quality 11 took 1,115 ms for 116.2 KiB. Level 9 cost 3.6 times as much CPU as level 6 for a body 3% smaller. Starlette's GZipMiddleware defaults to compresslevel=9.

Verify: a table of milliseconds and output size per level for your largest common response, so the level is chosen rather than inherited.

Compressing a 1.1 MB JSON response A grid of 5 rows by 4 columns. Compressing a 1.1 MB JSON response codec time output ratio gzip level 9 (Starlette default) 35.4 ms 192.5 KiB 5.9 gzip level 6 9.8 ms 198.0 KiB 5.7 gzip level 1 4.3 ms 226.7 KiB 5.0 brotli quality 4 5.7 ms 150.4 KiB 7.5 brotli quality 11 1,115 ms 116.2 KiB 9.7 Python 3.14 zlib, brotli 1.2.0; single-threaded timings.

2. Find where the middleware compresses

Starlette 1.7 compresses each body chunk on the event loop when it is smaller than thread_minimum_size, 128 KiB by default, and in a worker thread otherwise, under its own limit of 40 concurrent threads:

app = Starlette(routes=routes, middleware=[Middleware(GZipMiddleware)])
# defaults: minimum_size=500, compresslevel=9, thread_minimum_size=128 * 1024

Measured with a 117 KiB response: the middleware compressed every body on the loop, at about 3.4 ms each. Throughput was 291 responses per second with the server at 97% of one core, and the /ping probe — which does no work — had a median of 51.8 ms and a p99 of 104.6 ms, because each ping queued behind compressions. With an uncompressed 117 KiB response, the same probe's p99 was 2.0 ms. Older Starlette versions, and many hand-written middlewares, compress everything on the loop regardless of size; check the version you run.

Verify: time a trivial route while compressed responses are under load; if its p99 is several times its idle value, compression is on the loop.

3. Move compression off the loop for mid-sized bodies

zlib releases the GIL while it compresses, so threads give both responsiveness and parallelism. Lower the threshold so every compressed body goes to a thread:

app = Starlette(
    routes=routes,
    middleware=[Middleware(GZipMiddleware, compresslevel=6, thread_minimum_size=0)],
)

Measured with the 117 KiB response and thread_minimum_size=0 at level 9: 3,247 responses per second — 11 times the inline figure — with the probe at 0.3 ms median and 1.9 ms p99. The server process used 12.8 cores. Dropping to level 1 while staying inline gave 2,590 per second but a probe p99 of 12.4 ms: a cheaper level shortens the blocking, but does not remove it. Thread handoff costs some tens of microseconds per body, so for small JSON responses of a few KiB, inline compression at level 6 is cheaper than the handoff; minimum_size (500 bytes by default) skips the tiniest bodies entirely.

Verify: after the change, the probe's p99 under compressed load is within a few milliseconds of its idle value.

/ping p99 while serving 117 KiB gzip responses 4 horizontal bars comparing inline, level 9 (default below 128 KiB) with the others. /ping p99 while serving 117 KiB gzip responses inline, level 9 (default below 128 KiB) 104.6 ms inline, level 1 12.4 ms threads, level 9 (thread_minimum_size=0) 1.9 ms no compression 2.0 ms Throughput: 291, 2,590, 3,247 and 19,447 responses/s. Starlette 1.7.0, uvicorn 0.54, one worker process.

4. Lower the level and bound the CPU

Threads protect the loop, but the CPU is still spent — and with 40 compression threads allowed, one worker process can take most of the machine. Measured with the 1.1 MB response, which the default middleware already sends to threads:

# 1.1 MB response, compressed in threads in every case
#   level 9:   377 req/s, 192.5 KiB each, server at 1,504% CPU
#   level 6: 1,136 req/s, 198.0 KiB each, server at 1,483% CPU
#   level 1: 2,388 req/s, 226.7 KiB each, server at 1,312% CPU

At level 9, a single uvicorn worker used 15 cores to serve 377 responses per second; level 6 served three times as many on the same CPU for 3% more bytes. On a host running several workers or other services, that CPU comes from them. Set the level to 6 or lower, and if the host is shared, lower the thread limit or run compression in a reverse proxy sized for it. The probe's p99 stayed between 1.3 and 2.9 ms in all three runs — the loop was fine; the machine was busy.

Verify: server CPU per compressed response is known, and a full load test leaves headroom for the other processes on the host.

5. Compress once for responses that do not change

The cheapest compression is the one done ahead of time. For responses that are identical across requests — a configuration document, a catalogue page, a generated report — compress once and serve the bytes:

PRECOMPRESSED = gzip.compress(CATALOGUE_JSON, 9)        # at start-up or on change

async def catalogue(request):
    if "gzip" in request.headers.get("accept-encoding", ""):
        return Response(PRECOMPRESSED, media_type="application/json",
                        headers={"Content-Encoding": "gzip", "Vary": "Accept-Encoding"})
    return Response(CATALOGUE_JSON, media_type="application/json")

Measured: 15,952 responses per second at 198 KiB each, with the server under one core — 42 times the level 9 middleware's throughput. Precompressing can use the slowest, densest setting, since its cost is paid once. Always send Vary: Accept-Encoding, so caches do not serve gzip to a client that did not ask for it, and exclude these routes from the middleware so they are not compressed twice. Static assets are the same case; see serving static files from ASGI.

Verify: precompressed routes skip the middleware, send Content-Encoding and Vary, and decompress to the uncompressed body in a test.

Choosing where a response gets compressed A flow of 5 stages. Choosing where a response gets compressed Same bytes every time? precompress once, 15,952/s Under 500 bytes? send uncompressed A few KiB? inline, level 6 or lower Tens of KiB and up? threads, level 6 Shared host? bound threads or use the proxy Thresholds from Starlette 1.7's GZipMiddleware.

Verification

Response compression is configured well when:

  • The level is chosen from measurements, not left at 9: level 6 served 3 times the throughput for 3% more bytes.
  • No body tens of KiB in size is compressed on the event loop; a trivial route's p99 stays near idle under load.
  • Compression CPU is bounded relative to the host.
  • Unchanging responses are precompressed and served with Vary: Accept-Encoding.

Diagnostic Hook: when every route's latency rises together while one large-response route is busy, check the response sizes against the middleware's inline threshold. A 117 KiB response sat just below Starlette's 128 KiB thread_minimum_size and pushed /ping p99 to 104.6 ms.

Pitfalls & edge cases

  • The default level 9. Measured: 377 against 1,136 responses/s at level 6.
  • Bodies just under the thread threshold. Measured: probe p99 104.6 ms at 117 KiB.
  • Unbounded compression threads on a shared host. One worker used 15 cores.
  • Brotli at maximum quality on the fly. Measured: 1,115 ms per 1.1 MB body.

Frequently Asked Questions

Does Starlette's GZipMiddleware block the event loop?

For chunks under 128 KiB, yes: it compresses them inline. A 117 KiB response at level 9 raised a trivial route's p99 to 104.6 ms. Set thread_minimum_size lower to use threads.

What gzip level should an ASGI app use?

Level 6 or lower. Level 9 took 35.4 ms per 1.1 MB body against 9.8 ms at level 6, for an output only 3% smaller.

Does compressing in threads help with the GIL?

Yes for zlib, which releases the GIL while compressing. Moving 117 KiB compressions to threads raised throughput from 291 to 3,247 responses/s.

Should I precompress responses?

For responses that do not change between requests, yes. Precompressed bytes were served at 15,952 per second on under one core, against 377 with level 9 compression per request.