Choosing Between Uvicorn, Hypercorn and Granian¶
Uvicorn, Hypercorn and Granian all run any ASGI application, and benchmarks of "hello world" endpoints make them look wildly different. The differences are real — they come from the HTTP parser, the event loop and how much of the server is written in Python — and they mostly disappear once handlers do real work. Measured with wrk (2 threads, 64 connections) against one ASGI app on Python 3.14: on an endpoint that returns immediately, single-process throughput was 12,336 requests per second for Hypercorn, 19,100 for Uvicorn with asyncio and h11, 43,308 for Uvicorn with uvloop and httptools, and 87,990 for Granian; on an endpoint that awaits 10 ms, every server landed between 3,881 and 6,169 — close to the 6,400 per second that 64 connections × 10 ms allows. This guide shows how to read those numbers and what decides the choice once they stop mattering.
Prerequisites¶
- Python 3.11+,
pip install uvicorn[standard] hypercorn granian;wrkfor load (run here from a container). - ASGI basics, from ASGI Servers & Frameworks.
- Worker sizing, from sizing uvicorn workers for async services.
1. Benchmark with the same app and the same load¶
Compare servers on one application with two endpoints: one that does nothing, which measures server overhead, and one that awaits, which resembles real handlers:
# asgiapp.py — raw ASGI, so no framework overhead muddies the comparison
import asyncio
async def app(scope, receive, send):
if scope["type"] == "lifespan":
while True:
msg = await receive()
if msg["type"] == "lifespan.startup":
await send({"type": "lifespan.startup.complete"})
elif msg["type"] == "lifespan.shutdown":
await send({"type": "lifespan.shutdown.complete"})
return
if scope["path"] == "/io":
await asyncio.sleep(0.01) # stand-in for a 10 ms database call
body = b'{"ok":true}'
await send({"type": "http.response.start", "status": 200,
"headers": [(b"content-type", b"application/json"),
(b"content-length", str(len(body)).encode())]})
await send({"type": "http.response.body", "body": body})
uvicorn asgiapp:app --port 8000 --loop uvloop --http httptools --no-access-log
hypercorn asgiapp:app -b 127.0.0.1:8000
granian --interface asgi --port 8000 asgiapp:app
wrk -t2 -c64 -d6s http://127.0.0.1:8000/ # then /io
Run the load generator on separate cores (or a separate machine) and keep access logging off on every server — logging each request is often the most expensive thing a fast server does. Raw ASGI removes framework cost from the comparison; your framework adds the same overhead on every server.
Verify: each server returns identical responses for both endpoints before you trust any throughput number.
2. Read the I/O-bound numbers¶
The /io endpoint is closer to production. With 64 connections and a 10 ms await per request, no server can exceed 6,400 requests per second, and all of them came close:
| Server | / req/s |
/io req/s |
/io avg latency |
|---|---|---|---|
| Hypercorn | 12,336 | 3,881 | 16.4 ms |
| Uvicorn, asyncio + h11 | 19,100 | 4,367 | 14.6 ms |
| Uvicorn, uvloop + httptools | 43,308 | 5,366 | 11.9 ms |
| Granian | 87,990 | 5,975 | 10.7 ms |
| Uvicorn, 4 workers | 116,849 | 5,984 | 10.7 ms |
| Granian, 4 workers | 159,001 | 6,169 | 10.3 ms |
The spread shrinks from sevenfold to 1.6× because the 10 ms await dominates. What remains is per-request overhead added to latency: Hypercorn added about 6 ms over the 10 ms floor, Granian under 1 ms. If your handlers take 50–200 ms — typical for anything calling a database and an API — server overhead is a single-digit percentage of latency, and the choice should rest on the factors in step 4.
Verify: compute your handlers' median latency; if it is many times the server overhead you measured, throughput will not decide your choice.
3. Get the cheap Uvicorn speed-up¶
Uvicorn's own default stack matters. With uvicorn[standard] installed it auto-selects uvloop and httptools; with plain pip install uvicorn it runs the pure-Python h11 parser on the asyncio loop. Measured: 19,100 versus 43,308 requests per second on the empty endpoint, and 14.6 versus 11.9 ms on the I/O endpoint. Install the extra and make the choice explicit:
pip install "uvicorn[standard]"
uvicorn app:app --loop uvloop --http httptools --workers 4 --no-access-log
uvloop does not support Windows, and httptools is a C extension; on platforms without wheels, --loop asyncio --http h11 is the fallback. The loop comparison in isolation is in benchmarking uvloop against the default event loop.
Verify: the startup log names the loop and HTTP implementation in use; check it in production, not just locally.
4. Decide on protocol, operations and ecosystem¶
With throughput mostly settled by your own handlers, the deciding questions are practical:
- HTTP/2 or HTTP/3 to the server? Uvicorn speaks HTTP/1.1 and WebSockets only. Hypercorn supports HTTP/2 and HTTP/3; Granian supports HTTP/2. Most deployments terminate TLS and HTTP/2 at a load balancer and talk HTTP/1.1 to the app, which makes this moot — check your topology.
- trio? Hypercorn runs on trio as well as asyncio; the others are asyncio-only. Relevant for AnyIO-based apps that target trio, as in writing backend-agnostic code with AnyIO.
- Operational familiarity. Uvicorn is the most widely deployed and documented with FastAPI and Starlette; its behaviour under signals, reloads and gunicorn is well understood. Granian is newer, fast, and moves quickly — pin versions and read changelogs.
- Very cheap handlers at very high rates. Health-check fleets, proxies, routing layers: here server overhead is the workload, and Granian's measured lead is decisive.
def choose_server(needs_http3: bool, needs_trio: bool, handler_ms: float, rps_per_core: float) -> str:
if needs_http3 or needs_trio:
return "hypercorn"
if handler_ms < 2 and rps_per_core > 20_000:
return "granian" # server overhead dominates
return "uvicorn[standard]" # well-trodden default
Verify: the choice is written down with the deciding reason, and revisited if handler latency or protocol needs change.
5. Scale out with workers the same way on all three¶
All three run multiple worker processes; that, not the server, is how an asyncio service uses more than one core. Measured on the empty endpoint, four Uvicorn workers gave 116,849 requests per second and four Granian workers 159,001 — each about 2.7× and 1.8× its single-process figure on this shared machine, with wrk competing for the same cores. Size workers from CPU and memory, as described in sizing uvicorn workers for async services, and keep startup and shutdown in the ASGI lifespan so every server runs it the same way:
uvicorn app:app --workers 4 --loop uvloop --http httptools
granian --interface asgi --workers 4 app:app
hypercorn app:app --workers 4
Graceful shutdown semantics differ in details — how long each waits for in-flight requests, how it handles a second signal — so test a rolling deploy under load with the server you choose, using the checks in draining in-flight requests before shutdown.
Verify: a rolling restart under load drops no requests with your chosen server and worker count.
Verification¶
The server choice is sound when:
- It was benchmarked with your framework and a realistic endpoint, not only "hello world".
- Uvicorn, if chosen, runs uvloop and httptools in production, confirmed in startup logs.
- Protocol needs (HTTP/2, HTTP/3, trio) are met by the server or terminated in front of it.
- Rolling restarts under load lose no requests.
Diagnostic Hook: export request latency with and without the time spent inside your application (middleware can time the app; the difference from the load balancer's timing is server plus network). If server overhead is a large share of total latency, the server choice matters for you; if it is a few percent, optimise handlers instead.
Pitfalls & edge cases¶
- Benchmarking only empty endpoints. Real handlers shrink the differences to a few percent.
- Plain
pip install uvicorn. Runs the pure-Python stack, about half the throughput of[standard]on cheap endpoints. - Access logs during benchmarks. Logging can dominate a fast server's cost.
- Load generator on the same cores. Shared CPU understates every server's peak.
Frequently Asked Questions¶
Which is faster, Uvicorn, Hypercorn or Granian?
On an endpoint that returns immediately, Granian was fastest in testing at about 88,000 requests per second per process, Uvicorn with uvloop and httptools about 43,000, and Hypercorn about 12,000. With a handler that awaits 10 ms, all three were within 1.6 times of each other.
Does uvicorn[standard] make a difference?
Yes: it installs uvloop and httptools. In testing, Uvicorn with them served 43,308 requests per second on an empty endpoint against 19,100 with asyncio and h11.
Which ASGI server supports HTTP/2 and HTTP/3?
Hypercorn supports HTTP/2 and HTTP/3; Granian supports HTTP/2; Uvicorn supports HTTP/1.1 and WebSockets. Many deployments terminate HTTP/2 at a load balancer anyway.
Should I switch ASGI servers for performance?
Only if server overhead is a significant share of your request latency. For handlers taking tens of milliseconds, the measured differences were a few milliseconds per request.
Related¶
- ASGI Servers & Frameworks — up to the topic overview.
- Writing pure ASGI middleware — the same raw interface used in this benchmark.
- Network I/O & Protocol Handling — the section overview.