Skip to content

Designing Sans-I/O Protocol Libraries in Python

Most protocol code is written twice by accident. The parser is tangled into a while True: data = await reader.read() loop, so supporting a blocking client means copying it, testing it means opening sockets, and supporting trio means copying it again. The sans-I/O design splits a protocol implementation into a pure state machine that consumes bytes and emits events, plus a thin driver per I/O model that moves bytes between the socket and the machine. It is how h11, h2 and wsproto work, and why httpx and hypercorn can share them. This guide builds a small framed protocol that way and drives it from asyncio, from a blocking socket, and from a test that feeds it one byte at a time.

Prerequisites

1. Define the boundary: bytes in, events out

A sans-I/O object has no socket, no loop and no awaits. Its whole interface is three operations: hand it bytes you received, ask it for the next complete event, and ask it to encode something you want to send. The example protocol frames SET <key> <length>\r\n<payload>\r\n:

from dataclasses import dataclass


@dataclass
class Message:
    key: str
    value: bytes


class ProtocolError(Exception):
    pass


class KVProtocol:
    MAX_LINE = 1024

    def __init__(self) -> None:
        self._buf = bytearray()
        self._pending: tuple[str, int] | None = None

    def receive_data(self, data: bytes) -> None:
        self._buf += data

    def next_event(self) -> Message | None:
        if self._pending is None:
            end = self._buf.find(b"\r\n")
            if end < 0:
                if len(self._buf) > self.MAX_LINE:
                    raise ProtocolError("header too long")
                return None                              # need more bytes
            parts = bytes(self._buf[:end]).split()
            if len(parts) != 3 or parts[0] != b"SET":
                raise ProtocolError(f"bad header {parts!r}")
            self._pending = (parts[1].decode(), int(parts[2]))
            del self._buf[:end + 2]
        key, n = self._pending
        if len(self._buf) < n + 2:
            return None
        value = bytes(self._buf[:n])
        del self._buf[:n + 2]
        self._pending = None
        return Message(key, value)

    @staticmethod
    def encode(msg: Message) -> bytes:
        return b"SET %s %d\r\n%s\r\n" % (msg.key.encode(), len(msg.value), msg.value)

next_event() returning None means "not enough bytes yet". It never blocks and never reads: the caller decides when more bytes arrive. The MAX_LINE check is the kind of limit that is easy to forget when parsing is buried in an I/O loop — without it, a peer that never sends \r\n grows the buffer without bound.

Verify: the class imports nothing from asyncio or socket.

Where each concern lives in a sans-I/O design 4 stacked layers. Where each concern lives in a sans-I/O design application handles Message events, decides what to send asyncio driver await reader.read(), writer.write(), drain() blocking driver sock.recv(), sock.sendall() protocol state machine receive_data(), next_event(), encode() Only the drivers touch sockets; the parser is written and tested once.

2. Write the asyncio driver

The driver is a loop that moves bytes and dispatches events, and nothing more:

import asyncio


async def serve(reader: asyncio.StreamReader, writer: asyncio.StreamWriter, store: dict) -> None:
    proto = KVProtocol()
    try:
        while data := await reader.read(65536):
            proto.receive_data(data)
            while (event := proto.next_event()) is not None:
                store[event.key] = event.value
    except ProtocolError as exc:
        log.warning("closing %s: %s", writer.get_extra_info("peername"), exc)
    finally:
        writer.close()
        await writer.wait_closed()


async def main() -> None:
    store: dict[str, bytes] = {}
    server = await asyncio.start_server(lambda r, w: serve(r, w, store), "127.0.0.1", 9000)
    async with server:
        await server.serve_forever()

The inner while drains every complete event from one read, because a single read can contain several frames — or less than one. Driving this server with a blocking client that sent 1,000 frames of growing size stored all 1,000 values intact.

Verify: send frames from a client that splits writes at random byte offsets; every value must arrive intact.

3. Write a blocking driver with the same parser

The point of the separation shows up the moment a second I/O model needs the protocol. A blocking client is a dozen lines because it reuses the parser unchanged:

import socket


class BlockingClient:
    def __init__(self, host: str, port: int) -> None:
        self._sock = socket.create_connection((host, port), timeout=5)
        self._proto = KVProtocol()

    def set(self, key: str, value: bytes) -> None:
        self._sock.sendall(KVProtocol.encode(Message(key, value)))

    def read_event(self) -> Message:
        while (event := self._proto.next_event()) is None:
            data = self._sock.recv(65536)
            if not data:
                raise ConnectionError("peer closed")
            self._proto.receive_data(data)
        return event

The same would apply to a trio or AnyIO driver, a Protocol-based driver using data_received, or a driver that reads from a file for replaying captured traffic. Every one of them gets the header-length limit and the framing rules for free, and none can diverge from the others. That is the practical alternative to the dual-API problem discussed in offering sync and async versions of one API.

Verify: a frame written by the blocking client is parsed by the asyncio server, and the reverse.

One parser, many drivers A flow of 4 stages. One parser, many drivers bytes from any source stream, socket, file, test receive_data(bytes) append to buffer next_event() Message or None driver dispatches per I/O model Adding an I/O model means writing a driver, never touching the parser.

4. Test the parser without sockets

Because the parser is pure, the hardest cases are cheap to test exhaustively. The case that breaks most hand-rolled protocol code is a frame split at every possible boundary, so feed the encoded stream one byte at a time:

def test_byte_at_a_time() -> None:
    proto = KVProtocol()
    stream = (KVProtocol.encode(Message("a", b"hello"))
              + KVProtocol.encode(Message("b", b"x\r\ny")))     # payload contains the delimiter
    events = []
    for b in stream:
        proto.receive_data(bytes([b]))
        while (ev := proto.next_event()) is not None:
            events.append(ev)
    assert events == [Message("a", b"hello"), Message("b", b"x\r\ny")]


def test_header_limit() -> None:
    proto = KVProtocol()
    proto.receive_data(b"S" * 2000)
    with pytest.raises(ProtocolError):
        proto.next_event()

The second message deliberately contains \r\n inside its payload; a parser that searched for the delimiter instead of counting payload bytes would split it. These tests run in microseconds, need no event loop, and are the natural target for property-based testing with Hypothesis: generate arbitrary messages, encode them, split the stream at arbitrary points, and assert the round trip.

Verify: the byte-at-a-time test passes, and mutating the parser to search for \r\n in the payload makes it fail.

What becomes testable once I/O is removed A grid of 4 rows by 3 columns. What becomes testable once I/O is removed test case parser inside I/O loop sans-I/O parser frame split at every byte needs a custom fake stream a for loop malformed or oversized input needs a socket peer one call run time per case milliseconds, real sockets microseconds deterministic timing-dependent always Splitting at every byte is the test that finds framing bugs, and it is trivial once the parser has no I/O.

5. Handle backpressure in the driver, not the parser

The parser is synchronous, so it cannot wait for anything — including a slow peer. Flow control belongs to the driver. On the sending side that means await writer.drain() after writes; on the receiving side it means not reading more than the application can process:

async def forward(proto: KVProtocol, reader, writer, out: asyncio.Queue) -> None:
    while data := await reader.read(65536):
        proto.receive_data(data)
        while (event := proto.next_event()) is not None:
            await out.put(event)          # bounded queue: stops reading when the app is behind

Because the read only happens after every event has been handed off, a full queue stops the driver reading, which fills the kernel's receive buffer, which makes the peer's sends block — backpressure all the way to the sender, as described in handling drain and write backpressure in asyncio streams.

The parser contributes one thing: bounded internal buffering. Limits like MAX_LINE, and a maximum declared payload length checked before waiting for the payload, keep a hostile peer from using the parser's buffer as unlimited storage.

Verify: with a consumer that stops reading from the queue, the server's memory stays flat while the client blocks.

Verification

A sans-I/O protocol is done when:

  • The protocol module imports no I/O library — no asyncio, socket, trio or ssl.
  • Two drivers exist and interoperate, proving the boundary is real.
  • Byte-at-a-time and malformed-input tests run without a loop or a socket.
  • Every buffer in the parser is bounded by an explicit limit.

Diagnostic Hook: count ProtocolError by reason at the driver and log the peer address; a burst from one address is a broken client or a probe, a slow rise across all peers usually means a protocol version mismatch after a deploy. Expose the parser's buffered byte count per connection as well — a connection whose buffer sits near the limit is a peer sending incomplete frames.

Pitfalls & edge cases

  • Raising inside the driver loop without closing the connection. After a protocol error the stream position is unknown; close, do not resynchronise.
  • Returning several events from one call as a list. It works, but a single next_event() loop is easier to reason about for backpressure.
  • Letting the parser own timeouts. Time is I/O; idle and read timeouts belong in the driver, as in adding read timeouts to asyncio streams.
  • Copying the buffer on every call. bytes(self._buf) on a large buffer is O(n) per event; slice only the parts you need.

Frequently Asked Questions

What does sans-I/O mean in Python?

A protocol implementation that does no I/O itself: it accepts bytes you received, returns parsed events, and encodes outgoing messages to bytes. Separate drivers for asyncio, threads or trio move the bytes. h11, h2 and wsproto are built this way.

Why write a sans-I/O protocol library?

The parser is written once and shared by every concurrency model, the hardest parsing cases can be tested without sockets in microseconds, and different drivers cannot drift apart in how they interpret the protocol.

Where does backpressure go in a sans-I/O design?

In the driver. The parser is synchronous and bounded; the driver decides when to read more and awaits drain when writing, so flow control follows each I/O model's own rules.

Is sans-I/O slower than parsing inside the read loop?

The extra method calls cost little compared with system calls and the parsing itself. The usual performance risk is unnecessary buffer copying, which is avoidable in either design.