Skip to content

Choosing Between Uvicorn, Hypercorn and Granian

Uvicorn, Hypercorn and Granian all run any ASGI application, and benchmarks of "hello world" endpoints make them look wildly different. The differences are real — they come from the HTTP parser, the event loop and how much of the server is written in Python — and they mostly disappear once handlers do real work. Measured with wrk (2 threads, 64 connections) against one ASGI app on Python 3.14: on an endpoint that returns immediately, single-process throughput was 12,336 requests per second for Hypercorn, 19,100 for Uvicorn with asyncio and h11, 43,308 for Uvicorn with uvloop and httptools, and 87,990 for Granian; on an endpoint that awaits 10 ms, every server landed between 3,881 and 6,169 — close to the 6,400 per second that 64 connections × 10 ms allows. This guide shows how to read those numbers and what decides the choice once they stop mattering.

Prerequisites

1. Benchmark with the same app and the same load

Compare servers on one application with two endpoints: one that does nothing, which measures server overhead, and one that awaits, which resembles real handlers:

# asgiapp.py — raw ASGI, so no framework overhead muddies the comparison
import asyncio


async def app(scope, receive, send):
    if scope["type"] == "lifespan":
        while True:
            msg = await receive()
            if msg["type"] == "lifespan.startup":
                await send({"type": "lifespan.startup.complete"})
            elif msg["type"] == "lifespan.shutdown":
                await send({"type": "lifespan.shutdown.complete"})
                return
    if scope["path"] == "/io":
        await asyncio.sleep(0.01)                      # stand-in for a 10 ms database call
    body = b'{"ok":true}'
    await send({"type": "http.response.start", "status": 200,
                "headers": [(b"content-type", b"application/json"),
                            (b"content-length", str(len(body)).encode())]})
    await send({"type": "http.response.body", "body": body})
uvicorn asgiapp:app --port 8000 --loop uvloop --http httptools --no-access-log
hypercorn asgiapp:app -b 127.0.0.1:8000
granian --interface asgi --port 8000 asgiapp:app
wrk -t2 -c64 -d6s http://127.0.0.1:8000/      # then /io

Run the load generator on separate cores (or a separate machine) and keep access logging off on every server — logging each request is often the most expensive thing a fast server does. Raw ASGI removes framework cost from the comparison; your framework adds the same overhead on every server.

Verify: each server returns identical responses for both endpoints before you trust any throughput number.

Single-process requests per second, endpoint that returns immediately 4 horizontal bars comparing Granian with the others. Single-process requests per second, endpoint that returns immediately Granian 87,990 req/s Uvicorn, uvloop + httptools 43,308 req/s Uvicorn, asyncio + h11 19,100 req/s Hypercorn 12,336 req/s wrk -t2 -c64 -d6s on the same host; Python 3.14; access logs off. On an empty endpoint the server's own overhead is all you measure, and it varies sevenfold.

2. Read the I/O-bound numbers

The /io endpoint is closer to production. With 64 connections and a 10 ms await per request, no server can exceed 6,400 requests per second, and all of them came close:

Server / req/s /io req/s /io avg latency
Hypercorn 12,336 3,881 16.4 ms
Uvicorn, asyncio + h11 19,100 4,367 14.6 ms
Uvicorn, uvloop + httptools 43,308 5,366 11.9 ms
Granian 87,990 5,975 10.7 ms
Uvicorn, 4 workers 116,849 5,984 10.7 ms
Granian, 4 workers 159,001 6,169 10.3 ms

The spread shrinks from sevenfold to 1.6× because the 10 ms await dominates. What remains is per-request overhead added to latency: Hypercorn added about 6 ms over the 10 ms floor, Granian under 1 ms. If your handlers take 50–200 ms — typical for anything calling a database and an API — server overhead is a single-digit percentage of latency, and the choice should rest on the factors in step 4.

Verify: compute your handlers' median latency; if it is many times the server overhead you measured, throughput will not decide your choice.

3. Get the cheap Uvicorn speed-up

Uvicorn's own default stack matters. With uvicorn[standard] installed it auto-selects uvloop and httptools; with plain pip install uvicorn it runs the pure-Python h11 parser on the asyncio loop. Measured: 19,100 versus 43,308 requests per second on the empty endpoint, and 14.6 versus 11.9 ms on the I/O endpoint. Install the extra and make the choice explicit:

pip install "uvicorn[standard]"
uvicorn app:app --loop uvloop --http httptools --workers 4 --no-access-log

uvloop does not support Windows, and httptools is a C extension; on platforms without wheels, --loop asyncio --http h11 is the fallback. The loop comparison in isolation is in benchmarking uvloop against the default event loop.

Verify: the startup log names the loop and HTTP implementation in use; check it in production, not just locally.

What differs beyond raw speed A grid of 5 rows by 4 columns. What differs beyond raw speed aspect Uvicorn Hypercorn Granian core Python + C parser pure Python Rust HTTP/2 no yes yes HTTP/3 no yes (aioquic) no trio support no yes no process manager --workers, or gunicorn --workers built-in workers Protocol and ecosystem needs usually narrow the choice before throughput does.

4. Decide on protocol, operations and ecosystem

With throughput mostly settled by your own handlers, the deciding questions are practical:

  • HTTP/2 or HTTP/3 to the server? Uvicorn speaks HTTP/1.1 and WebSockets only. Hypercorn supports HTTP/2 and HTTP/3; Granian supports HTTP/2. Most deployments terminate TLS and HTTP/2 at a load balancer and talk HTTP/1.1 to the app, which makes this moot — check your topology.
  • trio? Hypercorn runs on trio as well as asyncio; the others are asyncio-only. Relevant for AnyIO-based apps that target trio, as in writing backend-agnostic code with AnyIO.
  • Operational familiarity. Uvicorn is the most widely deployed and documented with FastAPI and Starlette; its behaviour under signals, reloads and gunicorn is well understood. Granian is newer, fast, and moves quickly — pin versions and read changelogs.
  • Very cheap handlers at very high rates. Health-check fleets, proxies, routing layers: here server overhead is the workload, and Granian's measured lead is decisive.
def choose_server(needs_http3: bool, needs_trio: bool, handler_ms: float, rps_per_core: float) -> str:
    if needs_http3 or needs_trio:
        return "hypercorn"
    if handler_ms < 2 and rps_per_core > 20_000:
        return "granian"                        # server overhead dominates
    return "uvicorn[standard]"                  # well-trodden default

Verify: the choice is written down with the deciding reason, and revisited if handler latency or protocol needs change.

5. Scale out with workers the same way on all three

All three run multiple worker processes; that, not the server, is how an asyncio service uses more than one core. Measured on the empty endpoint, four Uvicorn workers gave 116,849 requests per second and four Granian workers 159,001 — each about 2.7× and 1.8× its single-process figure on this shared machine, with wrk competing for the same cores. Size workers from CPU and memory, as described in sizing uvicorn workers for async services, and keep startup and shutdown in the ASGI lifespan so every server runs it the same way:

uvicorn app:app --workers 4 --loop uvloop --http httptools
granian --interface asgi --workers 4 app:app
hypercorn app:app --workers 4

Graceful shutdown semantics differ in details — how long each waits for in-flight requests, how it handles a second signal — so test a rolling deploy under load with the server you choose, using the checks in draining in-flight requests before shutdown.

Verify: a rolling restart under load drops no requests with your chosen server and worker count.

Which ASGI server for this service? A decision on What does the service need with 3 outcomes. Which ASGI server for this service? What does the service need? HTTP/3 or trio Hypercorn protocol coverage tiny handlers, huge rates Granian lowest overhead typical API handlers Uvicorn [standard] well-trodden default Once handlers take tens of milliseconds, the measured gaps shrink to a few percent.

Verification

The server choice is sound when:

  • It was benchmarked with your framework and a realistic endpoint, not only "hello world".
  • Uvicorn, if chosen, runs uvloop and httptools in production, confirmed in startup logs.
  • Protocol needs (HTTP/2, HTTP/3, trio) are met by the server or terminated in front of it.
  • Rolling restarts under load lose no requests.

Diagnostic Hook: export request latency with and without the time spent inside your application (middleware can time the app; the difference from the load balancer's timing is server plus network). If server overhead is a large share of total latency, the server choice matters for you; if it is a few percent, optimise handlers instead.

Pitfalls & edge cases

  • Benchmarking only empty endpoints. Real handlers shrink the differences to a few percent.
  • Plain pip install uvicorn. Runs the pure-Python stack, about half the throughput of [standard] on cheap endpoints.
  • Access logs during benchmarks. Logging can dominate a fast server's cost.
  • Load generator on the same cores. Shared CPU understates every server's peak.

Frequently Asked Questions

Which is faster, Uvicorn, Hypercorn or Granian?

On an endpoint that returns immediately, Granian was fastest in testing at about 88,000 requests per second per process, Uvicorn with uvloop and httptools about 43,000, and Hypercorn about 12,000. With a handler that awaits 10 ms, all three were within 1.6 times of each other.

Does uvicorn[standard] make a difference?

Yes: it installs uvloop and httptools. In testing, Uvicorn with them served 43,308 requests per second on an empty endpoint against 19,100 with asyncio and h11.

Which ASGI server supports HTTP/2 and HTTP/3?

Hypercorn supports HTTP/2 and HTTP/3; Granian supports HTTP/2; Uvicorn supports HTTP/1.1 and WebSockets. Many deployments terminate HTTP/2 at a load balancer anyway.

Should I switch ASGI servers for performance?

Only if server overhead is a significant share of your request latency. For handlers taking tens of milliseconds, the measured differences were a few milliseconds per request.