Writing a TCP Server with asyncio.start_server¶
asyncio.start_server turns a coroutine into a TCP server: for every accepted connection it calls your handler with a StreamReader and a StreamWriter, and each handler runs as its own task. That model scales to thousands of mostly idle connections cheaply — measured on Python 3.14, 5,000 connected echo clients and their 5,000 server-side handlers, all in one process, added 62 MiB of memory, and one round trip on every connection completed in 0.18 s. The model also has one trap that bites every first server: writer.write() never blocks, so a handler that writes without await writer.drain() buffers without limit. Writing 125 MiB to a client that was not reading grew the server's memory by 122 MiB without drain(), and by 0 MiB with it. This guide builds a framed, bounded, cleanly stoppable server.
Prerequisites¶
- Python 3.12+ (3.13 for
close_clients()), stdlib only. - Streams versus protocols, from Streams, Transports & Protocols.
- Cancellation, from Cancellation Patterns.
1. Write the connection handler¶
The handler owns one connection for its whole life. Read in a loop until the client closes, and always close the writer in finally:
import asyncio
async def handle(reader: asyncio.StreamReader, writer: asyncio.StreamWriter) -> None:
peer = writer.get_extra_info("peername")
try:
while line := await reader.readline(): # b"" when the client closes
response = process(line)
writer.write(response)
await writer.drain() # backpressure: see step 3
except ConnectionResetError:
pass # client vanished; nothing to answer
finally:
writer.close()
try:
await writer.wait_closed()
except ConnectionError:
pass
log.debug("closed %s", peer)
async def main() -> None:
server = await asyncio.start_server(handle, "0.0.0.0", 9000, backlog=1024, limit=64 * 1024)
async with server:
await server.serve_forever()
An exception escaping the handler is logged by asyncio and the connection is dropped, but other connections are unaffected — each handler is its own task. limit bounds how much readline or readuntil will buffer looking for a delimiter (default 64 KiB); a longer line raises LimitOverrunError or ValueError, which protects the server from a client that never sends a newline.
Verify: nc localhost 9000 gets responses to each line, and closing nc ends the handler without an error in the logs.
2. Frame messages explicitly¶
TCP is a byte stream: one write on the client can arrive as several reads, and several writes as one. A protocol needs framing. Line-delimited and length-prefixed are the two common forms:
import struct
MAX_FRAME = 1 << 20 # 1 MiB
async def read_frame(reader: asyncio.StreamReader) -> bytes | None:
try:
header = await reader.readexactly(4)
except asyncio.IncompleteReadError as exc:
if exc.partial:
raise ConnectionError("truncated header")
return None # clean close between frames
(length,) = struct.unpack("!I", header)
if length > MAX_FRAME:
raise ValueError(f"frame of {length} bytes exceeds limit")
return await reader.readexactly(length)
def write_frame(writer: asyncio.StreamWriter, payload: bytes) -> None:
writer.write(struct.pack("!I", len(payload)) + payload)
readexactly waits for exactly n bytes or raises IncompleteReadError with whatever partial data arrived, which distinguishes a clean close (no partial bytes) from a truncated frame. Always check the declared length against a limit before reading the body; otherwise a client can make the server allocate whatever size it claims. For protocols you design, a sans-I/O parser keeps framing testable without sockets, as in designing sans-I/O protocol libraries.
Verify: a client that sends a frame in single-byte writes, and one that sends ten frames in one write, both get correct responses.
3. Await drain() after writing¶
writer.write() copies data into the transport's buffer and returns immediately. If the peer reads slowly or not at all, that buffer grows. await writer.drain() pauses the handler while the buffer is above its high-water mark:
async def stream_to_client(writer: asyncio.StreamWriter, chunks) -> None:
for chunk in chunks:
writer.write(chunk)
await writer.drain() # waits while the kernel and transport buffers are full
Measured with a client that never read: writing 2,000 chunks of 64 KiB grew the server's resident memory by 122 MiB without drain(), and by 0 MiB with it — the handler simply waited. Without drain, a few slow clients receiving large responses can exhaust the server's memory. drain() also surfaces errors: if the connection was reset, it raises ConnectionResetError instead of letting writes disappear into a dead socket.
Verify: connect a client that stops reading during a large response; server memory stays flat and the handler is suspended in drain().
4. Bound connections and idle time¶
A server should not accept unlimited connections or keep idle ones forever. Count connections with a semaphore and time out reads:
MAX_CONNECTIONS = 10_000
slots = asyncio.Semaphore(MAX_CONNECTIONS)
async def handle(reader, writer) -> None:
if slots.locked():
writer.close() # at capacity: refuse quickly
return
async with slots:
try:
while True:
try:
async with asyncio.timeout(300): # idle timeout: 5 minutes
frame = await read_frame(reader)
except TimeoutError:
break
if frame is None:
break
write_frame(writer, process(frame))
await writer.drain()
finally:
writer.close()
Measured, 5,000 idle connections on both ends in one process took 62 MiB, so memory is rarely the first limit; file descriptors are (ulimit -n, often 1,024 by default — raise it for servers). The idle timeout reclaims connections from clients that vanished without closing, complementing TCP keepalive; the timeout patterns are covered in adding read timeouts to asyncio streams.
Verify: connecting MAX_CONNECTIONS + 1 clients gets the last one refused, and an idle client is disconnected after the timeout.
5. Stop gracefully¶
On shutdown, stop accepting, give connected clients time to finish, then close the rest:
import signal
async def main() -> None:
server = await asyncio.start_server(handle, "0.0.0.0", 9000)
stop = asyncio.Event()
asyncio.get_running_loop().add_signal_handler(signal.SIGTERM, stop.set)
async with server:
await stop.wait()
server.close() # no new connections
try:
async with asyncio.timeout(10):
await server.wait_closed() # 3.12+: waits for handlers to finish
except TimeoutError:
server.close_clients() # 3.13+: close remaining connections
await server.wait_closed()
From Python 3.12, wait_closed() waits for active connections to finish, not just for the listening socket to close, so it doubles as "drain". close_clients() (3.13) closes the stragglers, which raises in their handlers' reads and lets their finally blocks run. For protocols with sessions, tell clients you are going away — a goodbye frame — before closing, so they reconnect elsewhere instead of treating it as a failure. The service-level sequence is in Graceful Shutdown & Signals.
Verify: SIGTERM during a load test lets in-flight requests complete, and the process exits within the grace period.
Verification¶
A start_server-based server is sound when:
- Messages are framed explicitly with size limits on every read.
- Every write is followed by
await drain()in code that can send much data. - Connections and idle time are bounded, and file-descriptor limits are raised.
- Shutdown stops accepting, drains, then closes remaining clients.
Diagnostic Hook: export the number of open connections, bytes buffered in transports (writer.transport.get_write_buffer_size()), and handlers waiting in drain(). Write buffers growing without bound mean a code path writes without draining; many handlers parked in drain() mean clients are slower than the server's output — usually slow networks or stuck clients that the idle timeout should reclaim.
Pitfalls & edge cases¶
- Writing without drain(). Measured: 122 MiB buffered for one slow client.
- Assuming one write is one read. TCP has no message boundaries; frame explicitly.
- Trusting declared lengths. Check them against a limit before reading.
- Default file-descriptor limits. 1,024 descriptors caps the server long before memory does.
Frequently Asked Questions¶
How do I write a TCP server with asyncio?
Call asyncio.start_server(handler, host, port) with an async handler that takes a StreamReader and StreamWriter, reads framed messages in a loop, writes responses followed by await writer.drain(), and closes the writer in finally.
Why do I need await writer.drain() in asyncio?
write() only buffers. Without drain(), data for a slow client accumulates in memory: 122 MiB in testing for a client that did not read, versus nothing when the handler awaited drain().
How many connections can an asyncio TCP server handle?
Thousands per process for mostly idle connections: 5,000 clients and their handlers took 62 MiB in one process in testing. File-descriptor limits usually bind first.
How do I stop an asyncio server gracefully?
Call server.close() to stop accepting, await server.wait_closed() with a timeout (on 3.12+ it waits for handlers), then call server.close_clients() on 3.13+ for any that remain.
Related¶
- Streams, Transports & Protocols — up to the topic overview.
- Building a TCP proxy with asyncio streams — two connections per handler, with backpressure both ways.
- Network I/O & Protocol Handling — the section overview.