Modelling Connection Lifecycles as Async State Machines¶
Long-lived connections — to a broker, a WebSocket peer, a database replica — move through states: connecting, connected, backing off, closing, closed. Code that tracks those states with separate booleans (connecting, connected, closed) works in the common order of events and breaks in the rare ones, because an await in the middle of a transition lets another event arrive. The classic failure is the zombie: close() is called while a connect is in its handshake, the handshake completes afterwards, and the object is now both closed and connected, holding a socket nobody will ever close. To measure it, 10,000 random schedules of four concurrent connect, close and reconnect calls were run against two implementations on Python 3.14. The flag-based version ended closed but connected in 2,509 of them — a quarter. The version with one state variable and an explicit transition table ended in an invalid state 0 times; it discarded 4,553 connections that completed after a close won the race and rejected 31,123 events that were not valid in the current state. This guide models connection lifecycles so that late events cannot corrupt them.
Prerequisites¶
- Python 3.11+; standard library only.
- An actor owning the connection, from building an actor with an asyncio mailbox.
- Reconnection policy, from reconnecting WebSocket clients with backoff.
1. See how flags go wrong across an await¶
The flag-based connection checks its flags, starts the handshake, awaits it, and then sets connected:
class FlagConn:
def __init__(self) -> None:
self.connecting = self.connected = self.closed = False
async def connect(self) -> None:
if self.connecting or self.connected or self.closed:
return
self.connecting = True
self.sock = await open_socket() # TCP + TLS: milliseconds
self.connected, self.connecting = True, False
async def close(self) -> None:
self.closed = True
if self.connected:
self.sock.close()
self.connected = False
If close() runs while connect() is suspended in open_socket(), close sees connected == False and closes nothing; then connect resumes and sets connected = True on a closed object. In the fuzzed schedules — four operations, random start delays up to 3 ms, handshakes of 0–2 ms — this happened in 2,509 of 10,000 runs, and it was the only kind of invalid end state that occurred. Every individual method is reasonable; the bug is in the combinations, which is why it survives code review and unit tests that call one method at a time.
Verify: a test that interleaves close() with an in-progress connect() ends with no open socket.
2. Replace flags with one state and a transition table¶
Give the connection exactly one state, and list the transitions that are allowed. Anything not in the table is rejected:
import enum
class S(enum.Enum):
IDLE = "idle"
CONNECTING = "connecting"
CONNECTED = "connected"
BACKOFF = "backoff"
CLOSED = "closed"
TRANSITIONS: dict[tuple[S, str], S] = {
(S.IDLE, "connect"): S.CONNECTING,
(S.BACKOFF, "connect"): S.CONNECTING,
(S.CONNECTING, "established"): S.CONNECTED,
(S.CONNECTING, "failed"): S.BACKOFF,
(S.CONNECTED, "lost"): S.BACKOFF,
(S.IDLE, "close"): S.CLOSED,
(S.CONNECTING, "close"): S.CLOSED,
(S.CONNECTED, "close"): S.CLOSED,
(S.BACKOFF, "close"): S.CLOSED,
}
class Conn:
def __init__(self) -> None:
self.state = S.IDLE
def _transition(self, event: str) -> bool:
nxt = TRANSITIONS.get((self.state, event))
if nxt is None:
log.debug("ignored %s in state %s", event, self.state.value)
return False
log.info("%s --%s--> %s", self.state.value, event, nxt.value)
self.state = nxt
return True
Impossible combinations are now unrepresentable: there is no way to be "closed and connected" because there is one variable. CLOSED has no outgoing transitions in this table, so nothing can resurrect a closed connection by accident. Logging every transition gives a complete history of each connection for free, which is most of what is needed to debug a reconnect storm later.
Verify: every assignment to self.state happens inside _transition.
3. Re-check the state after every await¶
The table does not remove the await; it makes the post-await check explicit. The handshake's result is applied only if the state is still the one that started it:
async def connect(self) -> None:
if not self._transition("connect"):
return # already connecting, connected or closed
try:
sock = await open_socket()
except OSError:
self._transition("failed")
return
if not self._transition("established"): # close() won the race
sock.close() # do not keep a socket for a closed object
return
self.sock = sock
async def close(self) -> None:
was = self.state
if self._transition("close") and was is S.CONNECTED:
self.sock.close()
In the fuzzed runs this path ran 4,553 times: a handshake completed, found the state already CLOSED, and closed its new socket. That is the zombie from step 1, caught and cleaned up, in nearly half of the runs. The 31,123 rejected events were mostly duplicate connects and closes, which are harmless when rejected and dangerous when acted on twice.
Verify: after any sequence of connect, close and reconnect calls, the number of open sockets equals 1 if the state is CONNECTED and 0 otherwise — the invariant the fuzz test asserted.
4. Run the machine inside an actor¶
The state machine removes impossible states; putting it inside an actor removes concurrent transitions altogether. Events become messages, and the actor applies them one at a time:
class ConnActor:
def __init__(self) -> None:
self.conn = Conn()
self.mailbox: asyncio.Queue[str] = asyncio.Queue()
async def run(self) -> None:
while (event := await self.mailbox.get()) != "stop":
match event:
case "connect":
await self.conn.connect()
case "close":
await self.conn.close()
case "lost":
self.conn._transition("lost")
asyncio.get_running_loop().call_later(self.backoff(), self.mailbox.put_nowait, "connect")
Now close() cannot run during the handshake at all — it waits in the mailbox until connect() finishes — so the "close won the race" branch becomes defensive rather than load-bearing. The trade-off is responsiveness: a close requested during a 10-second handshake waits for it, which may call for cancelling the handshake task on a close message. Scheduling the reconnect as a future message rather than sleeping inside the handler keeps the actor responsive to close during backoff. Reconnection with jittered delays is covered in reconnecting WebSocket clients with backoff.
Verify: while the actor is in BACKOFF, a close message takes effect immediately and no reconnect follows.
5. Fuzz the lifecycle in tests¶
The flag bug in step 1 needed specific interleavings, which ordinary tests do not produce. A small randomized test does:
import random
async def trial(rng: random.Random) -> Conn:
c = Conn()
async def later(op) -> None:
await asyncio.sleep(rng.uniform(0, 0.003))
await op()
ops = [rng.choice([c.connect, c.close, c.reconnect]) for _ in range(4)]
await asyncio.gather(*(later(op) for op in ops))
return c
async def test_lifecycle_invariants() -> None:
rng = random.Random(7)
for _ in range(10_000):
c = await trial(rng)
assert c.open_sockets == (1 if c.state is S.CONNECTED else 0)
assert c.state is not S.CONNECTING
The measured implementation supported reconnect with one extra, explicit event — reopen, allowed from CONNECTED and CLOSED back to IDLE — so that leaving CLOSED is always a deliberate act rather than a side effect. The fixed seed makes failures reproducible; the invariant is the one property that must hold whatever happened. Against the flag-based class this test failed on 2,509 of its 10,000 trials. For deeper exploration, Hypothesis can generate the operation sequences and shrink failures to a minimal case, as in property-based testing of async code with Hypothesis.
Verify: the fuzz test fails when the post-await state check in connect is removed.
Verification¶
A connection lifecycle is modelled safely when:
- There is one state variable, changed only through a transition table.
- Every
awaitinside a transition is followed by a re-check that the state has not moved on. - Late results are cleaned up, such as a socket that finished connecting after a close.
- A fuzz test with random interleavings checks the invariants on every run.
Diagnostic Hook: log every transition with the connection's identifier and count rejected events per type. A rising count of rejected established events means closes are regularly racing handshakes — usually a shutdown path or a health check that closes connections too eagerly.
Pitfalls & edge cases¶
- Boolean flags for state. Measured: closed-but-connected in a quarter of fuzzed schedules.
- Setting state after an await without re-checking. That is where the zombie is created.
- Sleeping in the actor during backoff. It cannot react to
closeuntil the sleep ends. - Transitions out of
CLOSED. Make reopening an explicit new object or an explicit event.
Frequently Asked Questions¶
How do I implement a state machine for an asyncio connection?
Keep one enum state, define allowed (state, event) transitions in a table, change state only through a function that consults it, and re-check the state after every await before applying a result. In fuzz testing this produced no invalid states in 10,000 schedules.
Why does my asyncio client end up connected after close() was called?
close() ran while connect() was awaiting the handshake; when the handshake finished, connect() set connected without checking. A flag-based client did this in 2,509 of 10,000 random schedules.
Should a connection state machine run inside an actor?
It helps: messages are applied one at a time, so close cannot interleave with connect. Keep backoff as a scheduled message rather than a sleep so close takes effect during backoff.
How do I test async state machines for race conditions?
Run thousands of randomized schedules of concurrent operations with a fixed seed and assert invariants such as 'one open socket if and only if connected' after each one.
Related¶
- Actors & Supervision — up to the topic overview.
- Bounding actor mailboxes under load — what happens when events arrive faster than they are handled.
- Concurrent Execution & Worker Patterns — the section overview.