Skip to content

Request-Reply Messaging Between asyncio Actors

Fire-and-forget messages cover updates, but most code also needs answers from an actor: the current balance, whether a key exists, the result of a validation. The standard pattern puts a future inside the message; the caller awaits it and the actor completes it. It is cheap — and it has one failure mode that turns a crash into a silent, permanent hang. Measured on Python 3.14 against a key-value actor: an ask round trip cost 4.54 µs sequentially and 6.25 µs each with 50,000 concurrent asks, against 0.062 µs for awaiting a plain coroutine that reads a dict. When the actor crashed on one message with three more asks queued behind it, all four callers were still waiting after one second — and would have waited forever. With a crash handler that failed the in-flight message and the queued asks, all four got RuntimeError: actor stopped immediately. Asks sent after the crash went into a mailbox that nobody would ever read. This guide builds request-reply that always answers.

Prerequisites

1. Put a future in the message

The caller creates a future, sends it with the request, and awaits it; the actor sets its result:

import asyncio
from dataclasses import dataclass, field


def _new_future() -> asyncio.Future:
    return asyncio.get_running_loop().create_future()


@dataclass
class Get:
    key: str
    reply: asyncio.Future = field(default_factory=_new_future)


@dataclass
class Put:
    key: str
    value: int


class KV:
    def __init__(self) -> None:
        self._data: dict[str, int] = {}
        self.mailbox: asyncio.Queue[Get | Put | None] = asyncio.Queue(maxsize=1000)

    async def run(self) -> None:
        while (msg := await self.mailbox.get()) is not None:
            match msg:
                case Put(key, value):
                    self._data[key] = value
                case Get(key, reply):
                    if not reply.done():               # the caller may have given up
                        reply.set_result(self._data.get(key))

    async def ask(self, key: str) -> int | None:
        msg = Get(key)
        await self.mailbox.put(msg)
        return await msg.reply

Measured: 4.54 µs per sequential ask, 6.25 µs each when 50,000 were in flight at once. That is roughly seventy times the cost of reading a dict directly, and still small next to any I/O. The reply.done() check matters: a caller that timed out or was cancelled leaves its future cancelled, and calling set_result on it would raise InvalidStateError inside the actor — crashing it over a caller's decision. Other completion pitfalls are in avoiding InvalidStateError when setting future results.

Verify: a caller that cancels its ask does not crash the actor.

One ask, end to end A sequence of 6 messages between 3 participants. One ask, end to end caller mailbox KV actor fut = loop.create_future() put(Get(key, fut)) await fut get() -> Get if not fut.done(): fut.set_result(value) value (4.5 us round trip) The future is the reply address; the actor never needs to know who asked.

2. Fail pending asks when the actor dies

If the actor raises while handling a message, its task ends. The future in that message is never completed, nor are the futures in messages still queued behind it:

async def run(self) -> None:
    msg = None
    try:
        while (msg := await self.mailbox.get()) is not None:
            await self.handle(msg)
    except BaseException as exc:
        error = RuntimeError(f"actor stopped: {exc!r}")
        self.closed = True
        if isinstance(msg, Get) and not msg.reply.done():
            msg.reply.set_exception(error)            # the message being handled
        while not self.mailbox.empty():
            queued = self.mailbox.get_nowait()
            if isinstance(queued, Get) and not queued.reply.done():
                queued.reply.set_exception(error)     # everything behind it
        raise

Measured with four asks where the second triggered a bug: without the handler, the first got its answer and the other three — including the one that crashed the actor — were still waiting after one second, with nothing in the logs except an unretrieved task exception. Failing only the queued asks left the in-flight caller hanging; failing both gave all three an immediate RuntimeError. Catching BaseException covers cancellation too: an actor cancelled during shutdown fails its pending asks instead of stranding them.

Verify: inject an exception into the handler and assert that every concurrent ask completes, with a value or an error, within a short timeout.

3. Refuse asks to a closed actor

After a crash, the mailbox still accepts messages; nothing will ever read them. Measured: an ask sent after the crash waited until its own timeout fired. A closed flag turns that into an immediate error:

class ActorStopped(RuntimeError):
    pass


async def ask(self, key: str, timeout: float | None = 1.0) -> int | None:
    if self.closed:
        raise ActorStopped(self.name)
    msg = Get(key)
    async with asyncio.timeout(timeout):
        await self.mailbox.put(msg)                   # may wait if the mailbox is full
        return await msg.reply

The timeout now bounds both waits — for space in the mailbox and for the reply — so a stalled actor cannot hold callers indefinitely. When the actor is restarted by a supervisor, as in supervising actors with restart strategies, the supervisor creates a fresh mailbox and clears the flag, so asks resume against the new incarnation.

Verify: ask on a crashed actor raises ActorStopped without waiting for its timeout.

Four concurrent asks, the second one crashes the actor A grid of 4 rows by 5 columns. Four concurrent asks, the second one crashes the actor crash handling ask 1 ask 2 (crashed) asks 3-4 later asks none answered waiting after 1 s waiting after 1 s wait until timeout fail queued asks only answered waiting after 1 s RuntimeError wait until timeout fail in-flight and queued answered RuntimeError RuntimeError wait until timeout ...plus a closed flag answered RuntimeError RuntimeError ActorStopped at once Without explicit handling, a crashed actor converts every pending request into a hang.

4. Ask other actors without deadlocking

Actors that ask each other can deadlock: if A, while handling a message, awaits an answer from B, and B, while handling a message, awaits an answer from A, both wait forever, because each one's mailbox is only read by the task that is now blocked. Two rules prevent it:

# Rule 1: an actor never awaits another actor's reply inside its handler
# when the other actor might ask back. Send a message instead, and handle
# the answer as a new message later.
async def handle(self, msg: Transfer) -> None:
    await self.ledger.mailbox.put(Reserve(msg.amount, reply_to=self.mailbox))


async def handle_reserved(self, msg: Reserved) -> None:
    self._pending.pop(msg.transfer_id).complete()


# Rule 2: if an actor must await a reply, the dependency graph is a tree:
# accounts may ask the ledger; the ledger never asks accounts.

The first rule is the actor model's original form — replies are just more messages — and avoids blocking entirely, at the cost of keeping per-request state across messages. The second is pragmatic: keep awaited asks, but only "downward" in a fixed hierarchy. A timeout on every ask, as in step 3, turns any deadlock that slips through into an error after the timeout, which is far easier to diagnose than a frozen service. The same reasoning about lock ordering appears in avoiding deadlocks with nested asyncio locks.

Verify: draw the "who awaits whom" graph between actors; it has no cycles.

5. Keep the ask API typed and narrow

A thin, typed client hides the message classes and futures from callers, so the protocol can change without touching every call site:

class KVClient:
    def __init__(self, actor: KV, timeout: float = 1.0) -> None:
        self._actor, self._timeout = actor, timeout

    async def get(self, key: str) -> int | None:
        return await self._actor.ask(key, timeout=self._timeout)

    async def put(self, key: str, value: int) -> None:
        if self._actor.closed:
            raise ActorStopped(self._actor.name)
        await self._actor.mailbox.put(Put(key, value))     # tell: no reply expected

Separating "tell" (no reply, like put) from "ask" (awaits a reply, like get) keeps the cheap path cheap and makes the blocking path visible in code review. Typing the client's methods also lets the checker enforce what the actor returns, which a Future of unspecified type does not — the approach from defining async protocols for dependency injection applies directly.

Verify: no code outside the actor's module constructs message objects or touches reply futures.

Tell or ask? A decision on What does the sender need with 4 outcomes. Tell or ask? What does the sender need? nothing back tell: put a message cheapest, no future a value now ask: future + timeout ~4.5 us round trip a value from a peer that may ask back reply as a later message no blocking a value from a lower layer ask with timeout acyclic only Every ask carries a timeout, and every actor fails its pending asks when it stops.

Verification

Request-reply between actors is reliable when:

  • Every ask carries a future and a timeout, and the actor checks done() before completing it.
  • A crashing actor fails the in-flight message and every queued ask.
  • Asks to a closed actor fail immediately rather than waiting for a timeout.
  • The graph of actors awaiting each other is acyclic.

Diagnostic Hook: count asks that end in TimeoutError per actor. A sudden rise across all callers of one actor means the actor stalled or died without failing its pending asks; a steady trickle means its handler occasionally exceeds the timeout, which is a capacity problem.

Pitfalls & edge cases

  • No crash handling. Measured: three of four callers still waiting after one second.
  • Failing only queued asks. The caller whose message caused the crash keeps waiting.
  • Setting a result on a cancelled future. It raises InvalidStateError in the actor.
  • Mutual asks. Two actors awaiting each other's replies deadlock.

Frequently Asked Questions

How do I get a return value from an asyncio actor?

Put a future in the message, await it in the caller, and have the actor call set_result when it handles the message. A round trip cost about 4.5 µs in testing.

Why do my requests hang when an actor crashes?

The futures in the crashing message and in queued messages are never completed. Catch exceptions in the actor's loop and call set_exception on the in-flight future and on every queued one; then mark the actor closed so later asks fail immediately.

Can two asyncio actors deadlock?

Yes, if each awaits a reply from the other while handling a message. Send replies as separate messages instead, or keep awaited asks to a hierarchy with no cycles, and put a timeout on every ask.

What is the difference between tell and ask?

Tell sends a message and continues without waiting; ask sends a message with a future and awaits the reply. Use tell for updates and ask only when a value is needed.