Designing Sans-I/O Protocol Libraries in Python¶
Most protocol code is written twice by accident. The parser is tangled into a while True: data = await reader.read() loop, so supporting a blocking client means copying it, testing it means opening sockets, and supporting trio means copying it again. The sans-I/O design splits a protocol implementation into a pure state machine that consumes bytes and emits events, plus a thin driver per I/O model that moves bytes between the socket and the machine. It is how h11, h2 and wsproto work, and why httpx and hypercorn can share them. This guide builds a small framed protocol that way and drives it from asyncio, from a blocking socket, and from a test that feeds it one byte at a time.
Prerequisites¶
- Python 3.11+, stdlib only.
- Asyncio streams, from Streams, Transports & Protocols.
- Framing basics, from implementing a length-prefixed framing protocol.
1. Define the boundary: bytes in, events out¶
A sans-I/O object has no socket, no loop and no awaits. Its whole interface is three operations: hand it bytes you received, ask it for the next complete event, and ask it to encode something you want to send. The example protocol frames SET <key> <length>\r\n<payload>\r\n:
from dataclasses import dataclass
@dataclass
class Message:
key: str
value: bytes
class ProtocolError(Exception):
pass
class KVProtocol:
MAX_LINE = 1024
def __init__(self) -> None:
self._buf = bytearray()
self._pending: tuple[str, int] | None = None
def receive_data(self, data: bytes) -> None:
self._buf += data
def next_event(self) -> Message | None:
if self._pending is None:
end = self._buf.find(b"\r\n")
if end < 0:
if len(self._buf) > self.MAX_LINE:
raise ProtocolError("header too long")
return None # need more bytes
parts = bytes(self._buf[:end]).split()
if len(parts) != 3 or parts[0] != b"SET":
raise ProtocolError(f"bad header {parts!r}")
self._pending = (parts[1].decode(), int(parts[2]))
del self._buf[:end + 2]
key, n = self._pending
if len(self._buf) < n + 2:
return None
value = bytes(self._buf[:n])
del self._buf[:n + 2]
self._pending = None
return Message(key, value)
@staticmethod
def encode(msg: Message) -> bytes:
return b"SET %s %d\r\n%s\r\n" % (msg.key.encode(), len(msg.value), msg.value)
next_event() returning None means "not enough bytes yet". It never blocks and never reads: the caller decides when more bytes arrive. The MAX_LINE check is the kind of limit that is easy to forget when parsing is buried in an I/O loop — without it, a peer that never sends \r\n grows the buffer without bound.
Verify: the class imports nothing from asyncio or socket.
2. Write the asyncio driver¶
The driver is a loop that moves bytes and dispatches events, and nothing more:
import asyncio
async def serve(reader: asyncio.StreamReader, writer: asyncio.StreamWriter, store: dict) -> None:
proto = KVProtocol()
try:
while data := await reader.read(65536):
proto.receive_data(data)
while (event := proto.next_event()) is not None:
store[event.key] = event.value
except ProtocolError as exc:
log.warning("closing %s: %s", writer.get_extra_info("peername"), exc)
finally:
writer.close()
await writer.wait_closed()
async def main() -> None:
store: dict[str, bytes] = {}
server = await asyncio.start_server(lambda r, w: serve(r, w, store), "127.0.0.1", 9000)
async with server:
await server.serve_forever()
The inner while drains every complete event from one read, because a single read can contain several frames — or less than one. Driving this server with a blocking client that sent 1,000 frames of growing size stored all 1,000 values intact.
Verify: send frames from a client that splits writes at random byte offsets; every value must arrive intact.
3. Write a blocking driver with the same parser¶
The point of the separation shows up the moment a second I/O model needs the protocol. A blocking client is a dozen lines because it reuses the parser unchanged:
import socket
class BlockingClient:
def __init__(self, host: str, port: int) -> None:
self._sock = socket.create_connection((host, port), timeout=5)
self._proto = KVProtocol()
def set(self, key: str, value: bytes) -> None:
self._sock.sendall(KVProtocol.encode(Message(key, value)))
def read_event(self) -> Message:
while (event := self._proto.next_event()) is None:
data = self._sock.recv(65536)
if not data:
raise ConnectionError("peer closed")
self._proto.receive_data(data)
return event
The same would apply to a trio or AnyIO driver, a Protocol-based driver using data_received, or a driver that reads from a file for replaying captured traffic. Every one of them gets the header-length limit and the framing rules for free, and none can diverge from the others. That is the practical alternative to the dual-API problem discussed in offering sync and async versions of one API.
Verify: a frame written by the blocking client is parsed by the asyncio server, and the reverse.
4. Test the parser without sockets¶
Because the parser is pure, the hardest cases are cheap to test exhaustively. The case that breaks most hand-rolled protocol code is a frame split at every possible boundary, so feed the encoded stream one byte at a time:
def test_byte_at_a_time() -> None:
proto = KVProtocol()
stream = (KVProtocol.encode(Message("a", b"hello"))
+ KVProtocol.encode(Message("b", b"x\r\ny"))) # payload contains the delimiter
events = []
for b in stream:
proto.receive_data(bytes([b]))
while (ev := proto.next_event()) is not None:
events.append(ev)
assert events == [Message("a", b"hello"), Message("b", b"x\r\ny")]
def test_header_limit() -> None:
proto = KVProtocol()
proto.receive_data(b"S" * 2000)
with pytest.raises(ProtocolError):
proto.next_event()
The second message deliberately contains \r\n inside its payload; a parser that searched for the delimiter instead of counting payload bytes would split it. These tests run in microseconds, need no event loop, and are the natural target for property-based testing with Hypothesis: generate arbitrary messages, encode them, split the stream at arbitrary points, and assert the round trip.
Verify: the byte-at-a-time test passes, and mutating the parser to search for \r\n in the payload makes it fail.
5. Handle backpressure in the driver, not the parser¶
The parser is synchronous, so it cannot wait for anything — including a slow peer. Flow control belongs to the driver. On the sending side that means await writer.drain() after writes; on the receiving side it means not reading more than the application can process:
async def forward(proto: KVProtocol, reader, writer, out: asyncio.Queue) -> None:
while data := await reader.read(65536):
proto.receive_data(data)
while (event := proto.next_event()) is not None:
await out.put(event) # bounded queue: stops reading when the app is behind
Because the read only happens after every event has been handed off, a full queue stops the driver reading, which fills the kernel's receive buffer, which makes the peer's sends block — backpressure all the way to the sender, as described in handling drain and write backpressure in asyncio streams.
The parser contributes one thing: bounded internal buffering. Limits like MAX_LINE, and a maximum declared payload length checked before waiting for the payload, keep a hostile peer from using the parser's buffer as unlimited storage.
Verify: with a consumer that stops reading from the queue, the server's memory stays flat while the client blocks.
Verification¶
A sans-I/O protocol is done when:
- The protocol module imports no I/O library — no
asyncio,socket,trioorssl. - Two drivers exist and interoperate, proving the boundary is real.
- Byte-at-a-time and malformed-input tests run without a loop or a socket.
- Every buffer in the parser is bounded by an explicit limit.
Diagnostic Hook: count ProtocolError by reason at the driver and log the peer address; a burst from one address is a broken client or a probe, a slow rise across all peers usually means a protocol version mismatch after a deploy. Expose the parser's buffered byte count per connection as well — a connection whose buffer sits near the limit is a peer sending incomplete frames.
Pitfalls & edge cases¶
- Raising inside the driver loop without closing the connection. After a protocol error the stream position is unknown; close, do not resynchronise.
- Returning several events from one call as a list. It works, but a single
next_event()loop is easier to reason about for backpressure. - Letting the parser own timeouts. Time is I/O; idle and read timeouts belong in the driver, as in adding read timeouts to asyncio streams.
- Copying the buffer on every call.
bytes(self._buf)on a large buffer is O(n) per event; slice only the parts you need.
Frequently Asked Questions¶
What does sans-I/O mean in Python?
A protocol implementation that does no I/O itself: it accepts bytes you received, returns parsed events, and encodes outgoing messages to bytes. Separate drivers for asyncio, threads or trio move the bytes. h11, h2 and wsproto are built this way.
Why write a sans-I/O protocol library?
The parser is written once and shared by every concurrency model, the hardest parsing cases can be tested without sockets in microseconds, and different drivers cannot drift apart in how they interpret the protocol.
Where does backpressure go in a sans-I/O design?
In the driver. The parser is synchronous and bounded; the driver decides when to read more and awaits drain when writing, so flow control follows each I/O model's own rules.
Is sans-I/O slower than parsing inside the read loop?
The extra method calls cost little compared with system calls and the parsing itself. The usual performance risk is unnecessary buffer copying, which is avoidable in either design.
Related¶
- Coroutine Design Patterns — up to the topic overview.
- Reading line-delimited protocols with asyncio streams — the stream-based alternative for simple text protocols.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.