Skip to content

Building an Actor with an asyncio Mailbox

asyncio removes data races between threads, but not races between tasks: any await between reading shared state and writing it back lets another task interleave. The usual fix is an asyncio.Lock around the critical section. The actor pattern is the other fix — give the state to one task, and let everything else send it messages through a queue — and it has properties a lock does not: no task ever waits for another to release anything, the state's invariants live in one place, and the owner can be supervised, restarted and observed as a unit. Measured on Python 3.14 with 10,000 concurrent increments of a balance, each with one await between read and write: without protection the final balance was 1 — 9,999 updates lost; with an asyncio.Lock it was correct and took 0.083 s; with an actor fed by 10,000 sending tasks it was correct and took 0.040 s. Raw mailbox throughput was 2.87 million messages per second, about 348 ns per message. This guide builds an actor that is correct, stoppable and observable.

Prerequisites

1. See the race the actor removes

A read-modify-write with an await in the middle is not atomic, even on one thread:

import asyncio


class Account:
    def __init__(self) -> None:
        self.balance = 0


async def deposit(acct: Account, amount: int) -> None:
    b = acct.balance
    await audit_log(amount)          # any await: a DB write, a sleep(0), a log flush
    acct.balance = b + amount


acct = Account()
await asyncio.gather(*(deposit(acct, 1) for _ in range(10_000)))
print(acct.balance)                   # measured: 1

Every task read 0, suspended at the await, and wrote back 1. The result was not "a few lost updates under load" but almost all of them, because every task reached its await before any of them resumed. The same bug hides in caches, rate-limiter counters, connection registries — anything that reads, awaits and then writes. The longer and more variable the awaited operation, the more interleavings exist; how to safely share state between async tasks and threads catalogues the variants.

Verify: a test that runs the operation concurrently a few thousand times checks the final state, not just the absence of exceptions.

10,000 concurrent deposits of 1, final balance 3 horizontal bars comparing no protection with the others. 10,000 concurrent deposits of 1, final balance no protection 1 (9,999 lost) asyncio.Lock 10,000 in 0.083 s actor with a mailbox 10,000 in 0.040 s Python 3.14; each deposit awaited asyncio.sleep(0) between reading and writing the balance. An await between read and write is enough to lose nearly every update.

2. Give the state to one task

An actor is a task that owns some state and a mailbox. Other code never touches the state; it puts messages in the mailbox, and the actor processes them one at a time:

from dataclasses import dataclass


@dataclass(frozen=True)
class Deposit:
    amount: int


@dataclass(frozen=True)
class Withdraw:
    amount: int


class AccountActor:
    def __init__(self, maxsize: int = 1000) -> None:
        self._balance = 0
        self.mailbox: asyncio.Queue[Deposit | Withdraw | None] = asyncio.Queue(maxsize)

    async def run(self) -> None:
        while (msg := await self.mailbox.get()) is not None:
            match msg:
                case Deposit(amount):
                    b = self._balance
                    await audit_log(amount)          # safe: nothing else touches _balance
                    self._balance = b + amount
                case Withdraw(amount):
                    if amount <= self._balance:
                        self._balance -= amount
                    else:
                        log.warning("rejected overdraft of %d", amount)

Measured with 10,000 tasks each awaiting mailbox.put(Deposit(1)): final balance 10,000, in 0.040 s against 0.083 s for the lock. The actor was faster because a contended asyncio.Lock wakes waiters one at a time through futures, while the actor's loop simply takes the next message. The await inside the handler is now safe — no other code can observe or modify _balance while it is suspended — and the business rule (no overdrafts) is enforced in exactly one place. Immutable message types make the protocol explicit and keep senders from mutating a message after sending it.

Verify: grep shows no access to the actor's state attributes from outside its class.

3. Start and stop the actor deliberately

An actor is a long-lived task, so it needs the same lifecycle discipline as any background task: an owner that starts it, a way to stop it that lets it finish its work, and a reference so it is not garbage-collected mid-run:

class AccountActor:
    ...
    async def stop(self) -> None:
        await self.mailbox.put(None)           # a sentinel queued behind real work


async def main() -> None:
    actor = AccountActor()
    async with asyncio.TaskGroup() as tg:
        tg.create_task(actor.run(), name="account-actor")
        await serve_requests(actor)            # sends messages
        await actor.stop()                     # drains, then run() returns

A sentinel queued behind outstanding messages drains the mailbox before the actor exits, which is usually what shutdown wants; cancelling the task instead stops it immediately and drops queued messages. Python 3.13's Queue.shutdown() offers a built-in alternative, covered in shutting down queues with Queue.shutdown. Running the actor inside the TaskGroup that also runs its clients means a crash in either cancels the other, rather than leaving clients sending to a dead mailbox.

Verify: stopping the application processes every message that was accepted before the stop, and leaves no actor tasks running.

A deposit through the actor A sequence of 6 messages between 4 participants. A deposit through the actor HTTP handler mailbox account actor audit log put(Deposit(1)) get() -> Deposit(1) b = balance await audit_log(1) done balance = b + 1: nobody else could interleave The await is inside the only task that touches the balance.

4. Process messages in batches when it pays

One message at a time is the simplest discipline, but an actor can drain whatever has accumulated and handle it together — a single database write for many deposits, say — without giving up exclusive ownership:

async def run(self) -> None:
    while (first := await self.mailbox.get()) is not None:
        batch = [first]
        while len(batch) < 500 and not self.mailbox.empty():
            msg = self.mailbox.get_nowait()
            if msg is None:
                await self._apply(batch)
                return
            batch.append(msg)
        await self._apply(batch)                 # one await for many messages


async def _apply(self, batch: list[Deposit]) -> None:
    total = sum(m.amount for m in batch)
    await self.db.execute("UPDATE accounts SET balance = balance + $1", total)
    self._balance += total

Batching changes the cost model: the mailbox overhead (about 348 ns per message measured) is negligible either way, but an awaited I/O call per message is not. Under light load batches are size one and latency is unchanged; under heavy load they grow and throughput rises with them. The cap keeps a single batch from monopolizing the actor's time. The general technique is in batching queue items by size and time.

Verify: under a burst, the logged batch sizes grow above one and the number of database round trips falls.

5. Make the actor observable

An actor's health is visible in three numbers: mailbox depth, time messages wait in the mailbox, and handler duration. Record them inside the actor, where they are cheap and accurate:

import time


class Envelope:
    __slots__ = ("msg", "enqueued")

    def __init__(self, msg: object) -> None:
        self.msg, self.enqueued = msg, time.perf_counter()


class ObservedActor:
    async def run(self) -> None:
        while (env := await self.mailbox.get()) is not None:
            started = time.perf_counter()
            MAILBOX_WAIT.observe(started - env.enqueued)
            await self.handle(env.msg)
            HANDLE_TIME.observe(time.perf_counter() - started)
            MAILBOX_DEPTH.set(self.mailbox.qsize())

Wait time rising while handler time stays flat means the actor is saturated — more messages arrive than one task can process — and the fix is sharding across several actors or batching. Handler time rising means a dependency slowed. This is the same split between queue wait and service time described in measuring queue wait and service time separately, applied to one actor.

Verify: a load test that exceeds the actor's capacity shows mailbox wait climbing while handler time stays constant.

Lock or actor for this shared state? A decision on What does the shared state look like with 4 outcomes. Lock or actor for this shared state? What does the shared state look like? no await between read and write no lock needed single-threaded loop an await inside, simple update asyncio.Lock 0.083 s for 10,000 rules across operations, needs supervision an actor with a mailbox 0.040 s for 10,000 one actor saturated shard actors by key see the sharding guide Actors earn their keep when the state has rules and a lifecycle, not just a counter.

Verification

An actor is correct and operable when:

  • Only the actor's task reads or writes its state, and all interaction goes through messages.
  • Concurrent tests check final state, which was wrong for 9,999 of 10,000 updates without protection.
  • The actor runs under an owner that stops it with a sentinel and awaits it.
  • Mailbox depth, wait time and handler time are exported.

Diagnostic Hook: alert when mailbox wait p99 exceeds the latency budget of the requests that send to the actor. Because the actor processes messages serially, its wait time is added to every caller's latency; a saturated actor shows up there before it shows up anywhere else.

Pitfalls & edge cases

  • Reading actor state from outside. It reintroduces the race the actor exists to remove.
  • Unbounded mailboxes. Under overload they grow without limit; see bounding actor mailboxes.
  • Slow handlers. One task processes everything; an awaited call per message caps throughput.
  • Mutable messages. A sender that changes a message after sending races with the actor.

Frequently Asked Questions

What is the actor model in Python asyncio?

A task that owns some state exclusively and receives messages through an asyncio.Queue, processing them one at a time. Other code never touches the state directly, so awaits inside the handler cannot cause races.

Is an actor faster than an asyncio.Lock?

In testing with 10,000 concurrent updates, each with an await inside the critical section, the actor took 0.040 s and the lock 0.083 s. The bigger benefits are a single place for invariants and a unit that can be supervised.

Do I need locks in asyncio if it is single-threaded?

Yes, whenever an await separates reading state from writing it. Without protection, 10,000 concurrent increments with one await in between produced a final value of 1.

How do I stop an asyncio actor cleanly?

Put a sentinel such as None on its mailbox behind the outstanding messages and await the actor's task; it drains what was accepted and returns. Cancelling the task drops queued messages instead.