Skip to content

Writing a TCP Server with asyncio.start_server

asyncio.start_server turns a coroutine into a TCP server: for every accepted connection it calls your handler with a StreamReader and a StreamWriter, and each handler runs as its own task. That model scales to thousands of mostly idle connections cheaply — measured on Python 3.14, 5,000 connected echo clients and their 5,000 server-side handlers, all in one process, added 62 MiB of memory, and one round trip on every connection completed in 0.18 s. The model also has one trap that bites every first server: writer.write() never blocks, so a handler that writes without await writer.drain() buffers without limit. Writing 125 MiB to a client that was not reading grew the server's memory by 122 MiB without drain(), and by 0 MiB with it. This guide builds a framed, bounded, cleanly stoppable server.

Prerequisites

1. Write the connection handler

The handler owns one connection for its whole life. Read in a loop until the client closes, and always close the writer in finally:

import asyncio


async def handle(reader: asyncio.StreamReader, writer: asyncio.StreamWriter) -> None:
    peer = writer.get_extra_info("peername")
    try:
        while line := await reader.readline():         # b"" when the client closes
            response = process(line)
            writer.write(response)
            await writer.drain()                       # backpressure: see step 3
    except ConnectionResetError:
        pass                                           # client vanished; nothing to answer
    finally:
        writer.close()
        try:
            await writer.wait_closed()
        except ConnectionError:
            pass
        log.debug("closed %s", peer)


async def main() -> None:
    server = await asyncio.start_server(handle, "0.0.0.0", 9000, backlog=1024, limit=64 * 1024)
    async with server:
        await server.serve_forever()

An exception escaping the handler is logged by asyncio and the connection is dropped, but other connections are unaffected — each handler is its own task. limit bounds how much readline or readuntil will buffer looking for a delimiter (default 64 KiB); a longer line raises LimitOverrunError or ValueError, which protects the server from a client that never sends a newline.

Verify: nc localhost 9000 gets responses to each line, and closing nc ends the handler without an error in the logs.

2. Frame messages explicitly

TCP is a byte stream: one write on the client can arrive as several reads, and several writes as one. A protocol needs framing. Line-delimited and length-prefixed are the two common forms:

import struct

MAX_FRAME = 1 << 20                                     # 1 MiB


async def read_frame(reader: asyncio.StreamReader) -> bytes | None:
    try:
        header = await reader.readexactly(4)
    except asyncio.IncompleteReadError as exc:
        if exc.partial:
            raise ConnectionError("truncated header")
        return None                                     # clean close between frames
    (length,) = struct.unpack("!I", header)
    if length > MAX_FRAME:
        raise ValueError(f"frame of {length} bytes exceeds limit")
    return await reader.readexactly(length)


def write_frame(writer: asyncio.StreamWriter, payload: bytes) -> None:
    writer.write(struct.pack("!I", len(payload)) + payload)

readexactly waits for exactly n bytes or raises IncompleteReadError with whatever partial data arrived, which distinguishes a clean close (no partial bytes) from a truncated frame. Always check the declared length against a limit before reading the body; otherwise a client can make the server allocate whatever size it claims. For protocols you design, a sans-I/O parser keeps framing testable without sockets, as in designing sans-I/O protocol libraries.

Verify: a client that sends a frame in single-byte writes, and one that sends ten frames in one write, both get correct responses.

One connection's handler A flow of 5 stages. One connection's handler accept new task per connection read frame readexactly + limit process your logic write + drain() backpressure client closes finally: close() The handler is ordinary sequential code; concurrency comes from one task per connection.

3. Await drain() after writing

writer.write() copies data into the transport's buffer and returns immediately. If the peer reads slowly or not at all, that buffer grows. await writer.drain() pauses the handler while the buffer is above its high-water mark:

async def stream_to_client(writer: asyncio.StreamWriter, chunks) -> None:
    for chunk in chunks:
        writer.write(chunk)
        await writer.drain()            # waits while the kernel and transport buffers are full

Measured with a client that never read: writing 2,000 chunks of 64 KiB grew the server's resident memory by 122 MiB without drain(), and by 0 MiB with it — the handler simply waited. Without drain, a few slow clients receiving large responses can exhaust the server's memory. drain() also surfaces errors: if the connection was reset, it raises ConnectionResetError instead of letting writes disappear into a dead socket.

Verify: connect a client that stops reading during a large response; server memory stays flat and the handler is suspended in drain().

Server memory while a client refuses to read 2 horizontal bars comparing write() without drain() with the others. Server memory while a client refuses to read write() without drain() +122 MiB write() + await drain() +0 MiB Python 3.14; 2,000 writes of 64 KiB to a client with a 64 KiB receive buffer that never read. drain() is the only backpressure a stream writer has.

4. Bound connections and idle time

A server should not accept unlimited connections or keep idle ones forever. Count connections with a semaphore and time out reads:

MAX_CONNECTIONS = 10_000
slots = asyncio.Semaphore(MAX_CONNECTIONS)


async def handle(reader, writer) -> None:
    if slots.locked():
        writer.close()                                  # at capacity: refuse quickly
        return
    async with slots:
        try:
            while True:
                try:
                    async with asyncio.timeout(300):    # idle timeout: 5 minutes
                        frame = await read_frame(reader)
                except TimeoutError:
                    break
                if frame is None:
                    break
                write_frame(writer, process(frame))
                await writer.drain()
        finally:
            writer.close()

Measured, 5,000 idle connections on both ends in one process took 62 MiB, so memory is rarely the first limit; file descriptors are (ulimit -n, often 1,024 by default — raise it for servers). The idle timeout reclaims connections from clients that vanished without closing, complementing TCP keepalive; the timeout patterns are covered in adding read timeouts to asyncio streams.

Verify: connecting MAX_CONNECTIONS + 1 clients gets the last one refused, and an idle client is disconnected after the timeout.

5. Stop gracefully

On shutdown, stop accepting, give connected clients time to finish, then close the rest:

import signal


async def main() -> None:
    server = await asyncio.start_server(handle, "0.0.0.0", 9000)
    stop = asyncio.Event()
    asyncio.get_running_loop().add_signal_handler(signal.SIGTERM, stop.set)

    async with server:
        await stop.wait()
        server.close()                                  # no new connections
        try:
            async with asyncio.timeout(10):
                await server.wait_closed()              # 3.12+: waits for handlers to finish
        except TimeoutError:
            server.close_clients()                      # 3.13+: close remaining connections
            await server.wait_closed()

From Python 3.12, wait_closed() waits for active connections to finish, not just for the listening socket to close, so it doubles as "drain". close_clients() (3.13) closes the stragglers, which raises in their handlers' reads and lets their finally blocks run. For protocols with sessions, tell clients you are going away — a goodbye frame — before closing, so they reconnect elsewhere instead of treating it as a failure. The service-level sequence is in Graceful Shutdown & Signals.

Verify: SIGTERM during a load test lets in-flight requests complete, and the process exits within the grace period.

Is start_server the right foundation? A decision on What does the server speak with 4 outcomes. Is start_server the right foundation? What does the server speak? custom framed protocol start_server + streams this guide huge rates of tiny messages Protocol class fewer copies HTTP ASGI server not hand-written needs TLS ssl=context argument same handler Streams are the right default; drop lower only when measurement says so.

Verification

A start_server-based server is sound when:

  • Messages are framed explicitly with size limits on every read.
  • Every write is followed by await drain() in code that can send much data.
  • Connections and idle time are bounded, and file-descriptor limits are raised.
  • Shutdown stops accepting, drains, then closes remaining clients.

Diagnostic Hook: export the number of open connections, bytes buffered in transports (writer.transport.get_write_buffer_size()), and handlers waiting in drain(). Write buffers growing without bound mean a code path writes without draining; many handlers parked in drain() mean clients are slower than the server's output — usually slow networks or stuck clients that the idle timeout should reclaim.

Pitfalls & edge cases

  • Writing without drain(). Measured: 122 MiB buffered for one slow client.
  • Assuming one write is one read. TCP has no message boundaries; frame explicitly.
  • Trusting declared lengths. Check them against a limit before reading.
  • Default file-descriptor limits. 1,024 descriptors caps the server long before memory does.

Frequently Asked Questions

How do I write a TCP server with asyncio?

Call asyncio.start_server(handler, host, port) with an async handler that takes a StreamReader and StreamWriter, reads framed messages in a loop, writes responses followed by await writer.drain(), and closes the writer in finally.

Why do I need await writer.drain() in asyncio?

write() only buffers. Without drain(), data for a slow client accumulates in memory: 122 MiB in testing for a client that did not read, versus nothing when the handler awaited drain().

How many connections can an asyncio TCP server handle?

Thousands per process for mostly idle connections: 5,000 clients and their handlers took 62 MiB in one process in testing. File-descriptor limits usually bind first.

How do I stop an asyncio server gracefully?

Call server.close() to stop accepting, await server.wait_closed() with a timeout (on 3.12+ it waits for handlers), then call server.close_clients() on 3.13+ for any that remain.