Skip to content

Bounding Actor Mailboxes Under Load

An actor processes one message at a time, so its capacity is fixed by how long each message takes. When messages arrive faster than that, they wait in the mailbox — and an unbounded asyncio.Queue lets them wait forever, growing memory and latency until something gives. A bounded mailbox forces a choice about what happens instead: the sender waits, the newest message is dropped, or the oldest is. Measured on Python 3.14 with an actor that took about 1 ms per message and a producer offering 2,000 messages per second for 4 seconds, twice the actor's capacity: the unbounded mailbox reached 4,338 messages, messages waited 2.4 s at the median and 4.6 s at p99, and the run took 8.7 s to drain. With 200 slots and blocking put, the depth never passed 200, p99 wait was 220 ms, and the producer was slowed to the actor's pace — 3,893 messages offered instead of 8,000. Dropping the newest messages lost 4,110 with the same latency; dropping the oldest lost 4,128 but cut the median wait to 99 ms, because the messages that survived were the freshest. This guide picks the right overload behaviour for an actor.

Prerequisites

1. Measure what an unbounded mailbox does at overload

asyncio.Queue() with no maxsize never refuses a message. Under sustained overload, its depth grows at the difference between arrival and service rates:

mailbox: asyncio.Queue = asyncio.Queue()             # unbounded


async def producer() -> None:
    start = time.perf_counter()
    for i in range(8000):                            # 2,000/s for 4 s
        await asyncio.sleep(max(0.0, start + i / 2000 - time.perf_counter()))
        await mailbox.put((time.perf_counter(), payload))   # never waits

Measured: maximum depth 4,338, peak traced memory 4.9 MiB for messages of 1 KiB each, median wait 2.36 s and p99 4.63 s. Every message was eventually processed, but most of them arrived seconds after anyone cared — a stale cache invalidation, a price update superseded long ago, a reply to a request whose caller timed out. With larger messages or a longer overload, memory is the next thing to go. The service-time signal stays flat throughout, which is why saturation is easy to miss without measuring queue wait, as discussed in diagnosing unbounded queue memory growth.

Verify: run the actor at twice its measured capacity for a minute and record the mailbox depth; if it grows without bound, the mailbox is unbounded in practice.

2,000 messages/s offered to a 1,000/s actor for 4 s A grid of 4 rows by 6 columns. 2,000 messages/s offered to a 1,000/s actor for 4 s policy offered processed lost max depth wait p50 / p99 unbounded 8,000 8,000 0 4,338 2,362 / 4,627 ms 200 slots, block sender 3,893 3,893 0 200 217 / 220 ms 200 slots, drop newest 8,001 3,891 4,110 200 215 / 218 ms 200 slots, drop oldest 8,000 3,872 4,128 200 99 / 193 ms Peak memory was 4.9 MiB unbounded and 0.3 MiB for every bounded policy.

2. Block senders when every message matters

When messages must not be lost and senders can slow down — internal pipeline stages, writes that must all happen — a bounded mailbox with an awaited put pushes the overload back to the producer:

class Actor:
    def __init__(self, capacity: int = 200) -> None:
        self.mailbox: asyncio.Queue = asyncio.Queue(maxsize=capacity)

    async def tell(self, msg: object, timeout: float = 1.0) -> None:
        async with asyncio.timeout(timeout):         # do not wait forever for space
            await self.mailbox.put(msg)

Measured: depth capped at 200, p99 wait 220 ms, 0.3 MiB peak — and the producer offered only 3,893 messages in the same 4 seconds, because each put waited for space. The overload did not disappear; it moved to the producer, where it can be handled deliberately. The capacity sets the worst-case wait: 200 messages at about 1 ms each is about 200 ms, which is what was measured. Choose the capacity from the latency you can tolerate, not from memory. The timeout on put matters when the producer is itself serving requests: a request that cannot even enqueue its message within a second should fail fast rather than pile up behind the actor.

Verify: at overload, senders' put latency rises and the mailbox depth stays at its capacity.

3. Drop the newest when the work can be refused

When the producer cannot slow down — an inbound request stream, a sensor feed — and a message that cannot be handled promptly can be rejected, drop at the door and tell the sender:

class Rejected(Exception):
    pass


def offer(self, msg: object) -> None:
    try:
        self.mailbox.put_nowait(msg)
    except asyncio.QueueFull:
        DROPPED.labels(policy="newest").inc()
        raise Rejected("actor overloaded")           # e.g. map to HTTP 503 upstream

Measured: 4,110 of 8,001 messages refused, and the ones accepted waited no more than 218 ms at p99. This is load shedding at the actor boundary: the rejected caller learns immediately and can retry elsewhere or back off, rather than waiting seconds for a result it no longer needs. It is the right default for actors that sit behind request handlers, and it pairs with load shedding when the event loop is overloaded at the service level.

Verify: at overload, the count of rejected messages rises and accepted-message latency stays within budget.

What should a full mailbox do? A decision on What does each message represent with 4 outcomes. What should a full mailbox do? What does each message represent? work that must all happen block sender, with a timeout overload moves upstream a request that can be refused drop newest, tell the sender 503 upstream a state update, newer supersedes older drop oldest or coalesce by key freshest wins anything, with no bound unbounded queue 4.6 s p99 at 2x load Bound the mailbox from the latency you can tolerate: capacity x service time.

4. Drop the oldest, or coalesce, for state updates

When each message is a new version of some state — a price, a position, a config value — an old message is worthless once a newer one exists. Dropping the oldest keeps the mailbox full of the freshest data:

def offer_latest(self, msg: object) -> None:
    if self.mailbox.full():
        self.mailbox.get_nowait()                    # discard the stalest message
        DROPPED.labels(policy="oldest").inc()
    self.mailbox.put_nowait(msg)

Measured: 4,128 messages dropped, and the median wait of those processed fell to 99 ms from about 215 ms, because the surviving messages had spent less time queued. For updates keyed by an entity, coalescing does better still: keep a dict of the latest message per key plus a queue of keys, and the mailbox never holds two versions of the same thing:

class CoalescingMailbox:
    def __init__(self) -> None:
        self._latest: dict[str, object] = {}
        self._keys: asyncio.Queue[str] = asyncio.Queue()

    def put(self, key: str, msg: object) -> None:
        if key not in self._latest:
            self._keys.put_nowait(key)               # queue each key once
        self._latest[key] = msg                      # overwrite older versions

    async def get(self) -> object:
        key = await self._keys.get()
        return self._latest.pop(key)

The coalescing mailbox's size is bounded by the number of distinct keys, not the message rate, and the actor always processes the newest version of each.

Verify: under overload, the version of each entity that the actor processes is the newest one sent.

5. Alert on depth, wait and drops together

A bounded mailbox converts overload from a slow-motion outage into numbers you can alert on. Export three, per actor:

async def run(self) -> None:
    while True:
        enqueued, msg = await self.mailbox.get()
        MAILBOX_WAIT.observe(time.perf_counter() - enqueued)
        MAILBOX_DEPTH.set(self.mailbox.qsize())
        await self.handle(msg)

Depth pinned at capacity means sustained overload. Wait near capacity × service time confirms it and tells you the latency callers are paying. A rising drop or rejection count tells you how much work is being shed. When all three move together, the actor needs more capacity — faster handling, batching, or more actors through sharding state by key — not a bigger mailbox, which only raises the latency ceiling.

Verify: a load test at twice capacity triggers the depth and drop alerts within one evaluation interval.

p99 mailbox wait at 2x overload 4 horizontal bars comparing unbounded with the others. p99 mailbox wait at 2x overload unbounded 4,627 ms 200 slots, block sender 220 ms 200 slots, drop newest 218 ms 200 slots, drop oldest 193 ms Actor service time about 1 ms per message; 1 KiB messages; 4-second overload. Every bounded policy held wait near capacity x service time; the unbounded one did not hold it at all.

Verification

An actor's mailbox is safe under load when:

  • The mailbox has a capacity chosen as tolerable latency divided by service time.
  • The overload policy matches the message type: block for must-process work, reject for refusable requests, drop oldest or coalesce for state.
  • Blocking put calls have a timeout, so senders cannot pile up indefinitely.
  • Depth, wait and drops are exported and alerted on together.

Diagnostic Hook: compare mailbox wait p99 with capacity × median service time. If wait is well below that, the actor has headroom; if it sits at that value, the actor is saturated and the mailbox is doing its job; if it is far above it, the mailbox is not actually bounded, or service time has a heavy tail worth investigating.

Pitfalls & edge cases

  • Unbounded mailboxes. Measured: 4,338 queued and 4.6 s p99 wait at 2x load.
  • Huge capacities. They behave like unbounded ones with extra steps; latency rises to match.
  • Blocking put without a timeout. Overload moves upstream and then has nowhere to stop.
  • Dropping silently. Count every drop, or overload looks like success.

Frequently Asked Questions

Should an asyncio actor's mailbox be bounded?

Yes. At twice the actor's capacity, an unbounded mailbox grew to 4,338 messages with a 4.6 s p99 wait; a 200-slot mailbox held p99 wait near 0.2 s under every overload policy.

What should happen when an actor's mailbox is full?

Block the sender (with a timeout) when every message must be processed, reject the newest when callers can be told no, and drop the oldest or coalesce by key when newer messages supersede older ones.

How big should an actor's mailbox be?

Capacity times service time is the worst-case wait, so divide the latency you can tolerate by the per-message handling time: 200 slots at about 1 ms per message gave the measured 220 ms p99 wait.

What is a coalescing mailbox?

A mailbox that keeps only the latest message per key, using a dict of latest values plus a queue of keys; its size is bounded by the number of distinct keys and the actor always sees the newest version.