Skip to content

Modelling Connection Lifecycles as Async State Machines

Long-lived connections — to a broker, a WebSocket peer, a database replica — move through states: connecting, connected, backing off, closing, closed. Code that tracks those states with separate booleans (connecting, connected, closed) works in the common order of events and breaks in the rare ones, because an await in the middle of a transition lets another event arrive. The classic failure is the zombie: close() is called while a connect is in its handshake, the handshake completes afterwards, and the object is now both closed and connected, holding a socket nobody will ever close. To measure it, 10,000 random schedules of four concurrent connect, close and reconnect calls were run against two implementations on Python 3.14. The flag-based version ended closed but connected in 2,509 of them — a quarter. The version with one state variable and an explicit transition table ended in an invalid state 0 times; it discarded 4,553 connections that completed after a close won the race and rejected 31,123 events that were not valid in the current state. This guide models connection lifecycles so that late events cannot corrupt them.

Prerequisites

1. See how flags go wrong across an await

The flag-based connection checks its flags, starts the handshake, awaits it, and then sets connected:

class FlagConn:
    def __init__(self) -> None:
        self.connecting = self.connected = self.closed = False

    async def connect(self) -> None:
        if self.connecting or self.connected or self.closed:
            return
        self.connecting = True
        self.sock = await open_socket()          # TCP + TLS: milliseconds
        self.connected, self.connecting = True, False

    async def close(self) -> None:
        self.closed = True
        if self.connected:
            self.sock.close()
        self.connected = False

If close() runs while connect() is suspended in open_socket(), close sees connected == False and closes nothing; then connect resumes and sets connected = True on a closed object. In the fuzzed schedules — four operations, random start delays up to 3 ms, handshakes of 0–2 ms — this happened in 2,509 of 10,000 runs, and it was the only kind of invalid end state that occurred. Every individual method is reasonable; the bug is in the combinations, which is why it survives code review and unit tests that call one method at a time.

Verify: a test that interleaves close() with an in-progress connect() ends with no open socket.

How a close() during the handshake leaves a zombie 3 lanes over time. How a close() during the handshake leaves a zombie connect() connecting = True await handshake (TCP + TLS) connected = True close() closed = True, nothing to close object state connecting closed + connecting closed AND connected time → Measured: 2,509 of 10,000 fuzzed schedules ended this way.

2. Replace flags with one state and a transition table

Give the connection exactly one state, and list the transitions that are allowed. Anything not in the table is rejected:

import enum


class S(enum.Enum):
    IDLE = "idle"
    CONNECTING = "connecting"
    CONNECTED = "connected"
    BACKOFF = "backoff"
    CLOSED = "closed"


TRANSITIONS: dict[tuple[S, str], S] = {
    (S.IDLE, "connect"): S.CONNECTING,
    (S.BACKOFF, "connect"): S.CONNECTING,
    (S.CONNECTING, "established"): S.CONNECTED,
    (S.CONNECTING, "failed"): S.BACKOFF,
    (S.CONNECTED, "lost"): S.BACKOFF,
    (S.IDLE, "close"): S.CLOSED,
    (S.CONNECTING, "close"): S.CLOSED,
    (S.CONNECTED, "close"): S.CLOSED,
    (S.BACKOFF, "close"): S.CLOSED,
}


class Conn:
    def __init__(self) -> None:
        self.state = S.IDLE

    def _transition(self, event: str) -> bool:
        nxt = TRANSITIONS.get((self.state, event))
        if nxt is None:
            log.debug("ignored %s in state %s", event, self.state.value)
            return False
        log.info("%s --%s--> %s", self.state.value, event, nxt.value)
        self.state = nxt
        return True

Impossible combinations are now unrepresentable: there is no way to be "closed and connected" because there is one variable. CLOSED has no outgoing transitions in this table, so nothing can resurrect a closed connection by accident. Logging every transition gives a complete history of each connection for free, which is most of what is needed to debug a reconnect storm later.

Verify: every assignment to self.state happens inside _transition.

3. Re-check the state after every await

The table does not remove the await; it makes the post-await check explicit. The handshake's result is applied only if the state is still the one that started it:

async def connect(self) -> None:
    if not self._transition("connect"):
        return                                   # already connecting, connected or closed
    try:
        sock = await open_socket()
    except OSError:
        self._transition("failed")
        return
    if not self._transition("established"):      # close() won the race
        sock.close()                             # do not keep a socket for a closed object
        return
    self.sock = sock


async def close(self) -> None:
    was = self.state
    if self._transition("close") and was is S.CONNECTED:
        self.sock.close()

In the fuzzed runs this path ran 4,553 times: a handshake completed, found the state already CLOSED, and closed its new socket. That is the zombie from step 1, caught and cleaned up, in nearly half of the runs. The 31,123 rejected events were mostly duplicate connects and closes, which are harmless when rejected and dangerous when acted on twice.

Verify: after any sequence of connect, close and reconnect calls, the number of open sockets equals 1 if the state is CONNECTED and 0 otherwise — the invariant the fuzz test asserted.

Connection states and the events that move between them A flow of 5 stages. Connection states and the events that move between them IDLE connect -> CONNECTING established / failed CONNECTED lost -> BACKOFF BACKOFF delay, then connect CLOSED from any state, final Nine allowed transitions; every other (state, event) pair is rejected and logged.

4. Run the machine inside an actor

The state machine removes impossible states; putting it inside an actor removes concurrent transitions altogether. Events become messages, and the actor applies them one at a time:

class ConnActor:
    def __init__(self) -> None:
        self.conn = Conn()
        self.mailbox: asyncio.Queue[str] = asyncio.Queue()

    async def run(self) -> None:
        while (event := await self.mailbox.get()) != "stop":
            match event:
                case "connect":
                    await self.conn.connect()
                case "close":
                    await self.conn.close()
                case "lost":
                    self.conn._transition("lost")
                    asyncio.get_running_loop().call_later(self.backoff(), self.mailbox.put_nowait, "connect")

Now close() cannot run during the handshake at all — it waits in the mailbox until connect() finishes — so the "close won the race" branch becomes defensive rather than load-bearing. The trade-off is responsiveness: a close requested during a 10-second handshake waits for it, which may call for cancelling the handshake task on a close message. Scheduling the reconnect as a future message rather than sleeping inside the handler keeps the actor responsive to close during backoff. Reconnection with jittered delays is covered in reconnecting WebSocket clients with backoff.

Verify: while the actor is in BACKOFF, a close message takes effect immediately and no reconnect follows.

5. Fuzz the lifecycle in tests

The flag bug in step 1 needed specific interleavings, which ordinary tests do not produce. A small randomized test does:

import random


async def trial(rng: random.Random) -> Conn:
    c = Conn()

    async def later(op) -> None:
        await asyncio.sleep(rng.uniform(0, 0.003))
        await op()

    ops = [rng.choice([c.connect, c.close, c.reconnect]) for _ in range(4)]
    await asyncio.gather(*(later(op) for op in ops))
    return c


async def test_lifecycle_invariants() -> None:
    rng = random.Random(7)
    for _ in range(10_000):
        c = await trial(rng)
        assert c.open_sockets == (1 if c.state is S.CONNECTED else 0)
        assert c.state is not S.CONNECTING

The measured implementation supported reconnect with one extra, explicit event — reopen, allowed from CONNECTED and CLOSED back to IDLE — so that leaving CLOSED is always a deliberate act rather than a side effect. The fixed seed makes failures reproducible; the invariant is the one property that must hold whatever happened. Against the flag-based class this test failed on 2,509 of its 10,000 trials. For deeper exploration, Hypothesis can generate the operation sequences and shrink failures to a minimal case, as in property-based testing of async code with Hypothesis.

Verify: the fuzz test fails when the post-await state check in connect is removed.

10,000 fuzzed schedules of connect, close and reconnect A grid of 2 rows by 4 columns. 10,000 fuzzed schedules of connect, close and reconnect implementation invalid end states late connections discarded events rejected three boolean flags 2,509 (closed but connected) 0 0 one state + transition table 0 4,553 31,123 Four concurrent operations per schedule, random delays up to 3 ms, handshakes of 0-2 ms.

Verification

A connection lifecycle is modelled safely when:

  • There is one state variable, changed only through a transition table.
  • Every await inside a transition is followed by a re-check that the state has not moved on.
  • Late results are cleaned up, such as a socket that finished connecting after a close.
  • A fuzz test with random interleavings checks the invariants on every run.

Diagnostic Hook: log every transition with the connection's identifier and count rejected events per type. A rising count of rejected established events means closes are regularly racing handshakes — usually a shutdown path or a health check that closes connections too eagerly.

Pitfalls & edge cases

  • Boolean flags for state. Measured: closed-but-connected in a quarter of fuzzed schedules.
  • Setting state after an await without re-checking. That is where the zombie is created.
  • Sleeping in the actor during backoff. It cannot react to close until the sleep ends.
  • Transitions out of CLOSED. Make reopening an explicit new object or an explicit event.

Frequently Asked Questions

How do I implement a state machine for an asyncio connection?

Keep one enum state, define allowed (state, event) transitions in a table, change state only through a function that consults it, and re-check the state after every await before applying a result. In fuzz testing this produced no invalid states in 10,000 schedules.

Why does my asyncio client end up connected after close() was called?

close() ran while connect() was awaiting the handshake; when the handshake finished, connect() set connected without checking. A flag-based client did this in 2,509 of 10,000 random schedules.

Should a connection state machine run inside an actor?

It helps: messages are applied one at a time, so close cannot interleave with connect. Keep backoff as a scheduled message rather than a sleep so close takes effect during backoff.

How do I test async state machines for race conditions?

Run thousands of randomized schedules of concurrent operations with a fixed seed and assert invariants such as 'one open socket if and only if connected' after each one.