Skip to content

Building a Minimal HTTP Server on asyncio Streams

Writing an HTTP server directly on asyncio.start_server is the best way to understand what an ASGI server does for you — and occasionally the right tool for an internal endpoint that needs nothing but a health check or a metrics page. The result is surprisingly fast. Measured with wrk (2 threads, 64 connections) on one core, a 40-line keep-alive HTTP/1.1 server on asyncio streams served 56,879 requests per second; with uvloop, 66,145. Uvicorn with uvloop and httptools, serving a raw ASGI app that returns the same bytes, did 38,674. The gap is the price of everything the minimal server does not do: chunked bodies, Expect: 100-continue, header validation, protocol upgrades, timeouts on every phase, graceful shutdown. This guide builds the minimal server, makes it robust against malformed and slow clients, and draws the line where you should switch to a real one.

Prerequisites

1. Accept connections and read the request head

asyncio.start_server calls your handler with a StreamReader and StreamWriter for each connection. An HTTP/1.1 request head ends with a blank line, so read until \r\n\r\n:

import asyncio

MAX_HEADER = 16 * 1024


async def handle(reader: asyncio.StreamReader, writer: asyncio.StreamWriter) -> None:
    try:
        while True:                                         # keep-alive: many requests per connection
            try:
                head = await asyncio.wait_for(reader.readuntil(b"\r\n\r\n"), timeout=15)
            except (asyncio.IncompleteReadError, TimeoutError, ConnectionError):
                return                                      # client closed or idled out
            except asyncio.LimitOverrunError:
                await respond(writer, 431, b"", keep=False)
                return
            if not await serve_one(head, reader, writer):
                return
    finally:
        writer.close()


async def main() -> None:
    server = await asyncio.start_server(handle, "127.0.0.1", 8080, limit=MAX_HEADER, backlog=1024)
    async with server:
        await server.serve_forever()

The limit argument caps how much readuntil buffers before raising LimitOverrunError, which bounds memory per connection — tested, a request with a 20 KB header got 431 Request Header Fields Too Large. The wait_for puts a deadline on receiving the whole head, so a client that sends one byte a second (a slowloris) is dropped after 15 s instead of holding the connection forever.

Verify: curl -v localhost:8080/ gets a response, and a request with an oversized header gets 431.

2. Parse the head and answer

The first line is method, target and version; each further line is a header. Parse defensively — a malformed request line must produce a 400, not an exception:

async def serve_one(head: bytes, reader, writer) -> bool:
    lines = head.decode("latin-1").split("\r\n")
    parts = lines[0].split(" ")
    if len(parts) != 3 or not parts[2].startswith("HTTP/1."):
        await respond(writer, 400, b"bad request\n", keep=False)
        return False
    method, target, version = parts
    headers = {}
    for line in lines[1:]:
        if line:
            name, sep, value = line.partition(":")
            if not sep:
                await respond(writer, 400, b"bad header\n", keep=False)
                return False
            headers[name.strip().lower()] = value.strip()
    if "transfer-encoding" in headers:                    # not supported: refuse, never guess
        await respond(writer, 501, b"", keep=False)
        return False
    length = int(headers.get("content-length", "0") or 0)
    body = await reader.readexactly(length) if length else b""
    keep = version == "HTTP/1.1" and headers.get("connection", "").lower() != "close"
    status, payload = route(method, target, body)
    await respond(writer, status, payload, keep)
    return keep


async def respond(writer, status: int, payload: bytes, keep: bool) -> None:
    reason = {200: "OK", 400: "Bad Request", 404: "Not Found", 431: "Request Header Fields Too Large",
              501: "Not Implemented"}.get(status, "")
    extra = "" if keep else "connection: close\r\n"
    head = f"HTTP/1.1 {status} {reason}\r\ncontent-length: {len(payload)}\r\n{extra}\r\n"
    writer.write(head.encode() + payload)
    await writer.drain()

Without the request-line check, a client sending garbage\r\n\r\n raised ValueError inside the connection callback — tested, asyncio logged "Unhandled exception in client_connected_cb" and the client got an empty response. Refusing Transfer-Encoding is a security decision: a server that half-understands chunked encoding behind a proxy that fully understands it is the recipe for request smuggling. Pipelined requests — two requests sent back to back without waiting — work automatically, because the loop reads the next head from the same buffer; tested, two pipelined GETs got two 200 responses.

Verify: send malformed request lines and headers with nc; each gets a 400 and the server logs no exceptions.

The per-connection loop of a minimal HTTP/1.1 server A flow of 5 stages. The per-connection loop of a minimal HTTP/1.1 server readuntil blank line limit + deadline parse, validate 400 / 431 / 501 read body Content-Length only route, respond drain() keep-alive? loop or close Every arrow is a place a hostile or broken client can stall or break the server.

3. Measure it against a real server

Throughput for trivial responses is where a minimal server looks best, because it skips work a real server must do:

python mini.py &                                       # stdlib event loop
wrk -t2 -c64 -d8s http://127.0.0.1:8080/              # 56,879 req/s

python mini.py uvloop &                                # same code on uvloop
wrk -t2 -c64 -d8s http://127.0.0.1:8080/              # 66,145 req/s

uvicorn hello_asgi:app --loop uvloop --http httptools  # raw ASGI app, same response
wrk -t2 -c64 -d8s http://127.0.0.1:8000/              # 38,674 req/s

Uvicorn's extra cost buys a full HTTP/1.1 parser with chunked bodies, ASGI message dispatch, flow control, keep-alive and request timeouts, lifespan, graceful shutdown and WebSocket upgrades. Any real application work per request — a database call, JSON serialization, a template — swamps the difference between 39k and 57k requests per second. Comparisons between real servers are in choosing between Uvicorn, Hypercorn and Granian.

Verify: benchmark your own handler on both; once it does real work, the gap shrinks to noise.

Requests per second for a 6-byte response 3 horizontal bars comparing minimal server, asyncio with the others. Requests per second for a 6-byte response minimal server, asyncio 56,879 req/s minimal server, uvloop 66,145 req/s Uvicorn + raw ASGI app 38,674 req/s wrk -t2 -c64 -d8s, one server process, same host. The minimal server is faster because it does far less.

4. Add graceful shutdown

A server that stops by being killed drops in-flight requests. Stop accepting, let open connections finish their current request, then exit:

import signal


async def main() -> None:
    server = await asyncio.start_server(handle, "127.0.0.1", 8080, limit=MAX_HEADER)
    stop = asyncio.Event()
    loop = asyncio.get_running_loop()
    for sig in (signal.SIGTERM, signal.SIGINT):
        loop.add_signal_handler(sig, stop.set)

    async with server:
        await stop.wait()
        server.close()                                  # stop accepting new connections
        try:
            async with asyncio.timeout(10):
                await server.wait_closed()              # 3.12+: waits for active connections
        except TimeoutError:
            server.close_clients()                      # 3.13+: force-close the stragglers

On Python 3.12 and later, Server.wait_closed() waits until all connection handlers have finished, which includes idle keep-alive connections waiting for their next request — the 15 s head timeout from step 1 bounds that. close_clients() (3.13+) closes whatever is left after the grace period. The full shutdown sequence for services is in Graceful Shutdown & Signals.

Verify: send SIGTERM during a load test; in-flight requests complete, new connections are refused, and the process exits within the grace period.

5. Know when to stop and use a real server

The minimal server is appropriate when the protocol surface is tiny and the clients are known: a health or metrics endpoint on an internal port, a test double for an HTTP dependency, an embedded control interface. It is the wrong choice for anything exposed to browsers or the internet.

# A fine use: a metrics endpoint alongside a non-HTTP service, no web framework needed
def route(method: str, target: str, body: bytes) -> tuple[int, bytes]:
    if method == "GET" and target == "/healthz":
        return 200, b"ok\n"
    if method == "GET" and target == "/metrics":
        return 200, render_metrics().encode()
    return 404, b"not found\n"

Everything this server lacks — chunked request bodies, 100-continue, HTTP/2, TLS configuration, header normalization, per-phase timeouts, protection against request smuggling — is a feature some client will eventually rely on or an attack someone will eventually try. For HTTP-facing applications, use an ASGI server, or aiohttp.web if you want a single library for client and server.

Verify: the minimal server listens only on an internal interface, and anything public goes through a real server.

Is a hand-written HTTP server appropriate here? A decision on Who are the clients with 3 outcomes. Is a hand-written HTTP server appropriate here? Who are the clients? orchestrator probes, scrapers minimal server ok internal port only tests faking a dependency minimal server ok full control of bytes browsers, public internet ASGI server or aiohttp.web full HTTP semantics Speed is not the reason to hand-roll HTTP; a tiny, known surface is.

Verification

A minimal HTTP server is acceptable when:

  • Header size and receive time are bounded for every request.
  • Malformed input produces 4xx responses, never unhandled exceptions.
  • Unsupported features are refused (Transfer-Encoding), not guessed at.
  • Shutdown is graceful and the server is not exposed publicly.

Diagnostic Hook: count connections, requests per connection and 4xx responses by status. A rising number of open connections with few requests means clients are idling or trickling — check the head timeout. 400s from a source you trust mean the client speaks more HTTP than your server does; that is the signal to move to a real server.

Pitfalls & edge cases

  • Parsing without validation. A malformed request line raised an unhandled exception in testing.
  • No read deadline. Slow clients hold connections indefinitely.
  • Half-supporting chunked encoding. Refuse it rather than risk request smuggling.
  • Benchmarking trivial responses. Real work erases the throughput advantage.

Frequently Asked Questions

How do I write an HTTP server with asyncio streams?

Use asyncio.start_server with a handler that reads the request head with readuntil(b"\r\n\r\n"), parses the request line and headers, reads a Content-Length body, writes the response and loops for keep-alive.

Is a hand-written asyncio HTTP server faster than Uvicorn?

For trivial responses, yes: 56,879 requests per second against 38,674 in testing. It does far less work per request, and real application work erases the difference.

How do I protect a minimal asyncio server from slow clients?

Put a deadline on reading the request head with asyncio.wait_for or asyncio.timeout, and cap buffered header size with start_server's limit argument.

When should I not write my own HTTP server?

For anything browser-facing or public. Use an ASGI server or aiohttp.web, which handle chunked bodies, protocol upgrades, timeouts and request smuggling defences.