Bounding Actor Mailboxes Under Load¶
An actor processes one message at a time, so its capacity is fixed by how long each message takes. When messages arrive faster than that, they wait in the mailbox — and an unbounded asyncio.Queue lets them wait forever, growing memory and latency until something gives. A bounded mailbox forces a choice about what happens instead: the sender waits, the newest message is dropped, or the oldest is. Measured on Python 3.14 with an actor that took about 1 ms per message and a producer offering 2,000 messages per second for 4 seconds, twice the actor's capacity: the unbounded mailbox reached 4,338 messages, messages waited 2.4 s at the median and 4.6 s at p99, and the run took 8.7 s to drain. With 200 slots and blocking put, the depth never passed 200, p99 wait was 220 ms, and the producer was slowed to the actor's pace — 3,893 messages offered instead of 8,000. Dropping the newest messages lost 4,110 with the same latency; dropping the oldest lost 4,128 but cut the median wait to 99 ms, because the messages that survived were the freshest. This guide picks the right overload behaviour for an actor.
Prerequisites¶
- Python 3.11+; standard library only.
- An actor, from building an actor with an asyncio mailbox.
- Queue backpressure basics, from bounded asyncio.Queue with backpressure.
1. Measure what an unbounded mailbox does at overload¶
asyncio.Queue() with no maxsize never refuses a message. Under sustained overload, its depth grows at the difference between arrival and service rates:
mailbox: asyncio.Queue = asyncio.Queue() # unbounded
async def producer() -> None:
start = time.perf_counter()
for i in range(8000): # 2,000/s for 4 s
await asyncio.sleep(max(0.0, start + i / 2000 - time.perf_counter()))
await mailbox.put((time.perf_counter(), payload)) # never waits
Measured: maximum depth 4,338, peak traced memory 4.9 MiB for messages of 1 KiB each, median wait 2.36 s and p99 4.63 s. Every message was eventually processed, but most of them arrived seconds after anyone cared — a stale cache invalidation, a price update superseded long ago, a reply to a request whose caller timed out. With larger messages or a longer overload, memory is the next thing to go. The service-time signal stays flat throughout, which is why saturation is easy to miss without measuring queue wait, as discussed in diagnosing unbounded queue memory growth.
Verify: run the actor at twice its measured capacity for a minute and record the mailbox depth; if it grows without bound, the mailbox is unbounded in practice.
2. Block senders when every message matters¶
When messages must not be lost and senders can slow down — internal pipeline stages, writes that must all happen — a bounded mailbox with an awaited put pushes the overload back to the producer:
class Actor:
def __init__(self, capacity: int = 200) -> None:
self.mailbox: asyncio.Queue = asyncio.Queue(maxsize=capacity)
async def tell(self, msg: object, timeout: float = 1.0) -> None:
async with asyncio.timeout(timeout): # do not wait forever for space
await self.mailbox.put(msg)
Measured: depth capped at 200, p99 wait 220 ms, 0.3 MiB peak — and the producer offered only 3,893 messages in the same 4 seconds, because each put waited for space. The overload did not disappear; it moved to the producer, where it can be handled deliberately. The capacity sets the worst-case wait: 200 messages at about 1 ms each is about 200 ms, which is what was measured. Choose the capacity from the latency you can tolerate, not from memory. The timeout on put matters when the producer is itself serving requests: a request that cannot even enqueue its message within a second should fail fast rather than pile up behind the actor.
Verify: at overload, senders' put latency rises and the mailbox depth stays at its capacity.
3. Drop the newest when the work can be refused¶
When the producer cannot slow down — an inbound request stream, a sensor feed — and a message that cannot be handled promptly can be rejected, drop at the door and tell the sender:
class Rejected(Exception):
pass
def offer(self, msg: object) -> None:
try:
self.mailbox.put_nowait(msg)
except asyncio.QueueFull:
DROPPED.labels(policy="newest").inc()
raise Rejected("actor overloaded") # e.g. map to HTTP 503 upstream
Measured: 4,110 of 8,001 messages refused, and the ones accepted waited no more than 218 ms at p99. This is load shedding at the actor boundary: the rejected caller learns immediately and can retry elsewhere or back off, rather than waiting seconds for a result it no longer needs. It is the right default for actors that sit behind request handlers, and it pairs with load shedding when the event loop is overloaded at the service level.
Verify: at overload, the count of rejected messages rises and accepted-message latency stays within budget.
4. Drop the oldest, or coalesce, for state updates¶
When each message is a new version of some state — a price, a position, a config value — an old message is worthless once a newer one exists. Dropping the oldest keeps the mailbox full of the freshest data:
def offer_latest(self, msg: object) -> None:
if self.mailbox.full():
self.mailbox.get_nowait() # discard the stalest message
DROPPED.labels(policy="oldest").inc()
self.mailbox.put_nowait(msg)
Measured: 4,128 messages dropped, and the median wait of those processed fell to 99 ms from about 215 ms, because the surviving messages had spent less time queued. For updates keyed by an entity, coalescing does better still: keep a dict of the latest message per key plus a queue of keys, and the mailbox never holds two versions of the same thing:
class CoalescingMailbox:
def __init__(self) -> None:
self._latest: dict[str, object] = {}
self._keys: asyncio.Queue[str] = asyncio.Queue()
def put(self, key: str, msg: object) -> None:
if key not in self._latest:
self._keys.put_nowait(key) # queue each key once
self._latest[key] = msg # overwrite older versions
async def get(self) -> object:
key = await self._keys.get()
return self._latest.pop(key)
The coalescing mailbox's size is bounded by the number of distinct keys, not the message rate, and the actor always processes the newest version of each.
Verify: under overload, the version of each entity that the actor processes is the newest one sent.
5. Alert on depth, wait and drops together¶
A bounded mailbox converts overload from a slow-motion outage into numbers you can alert on. Export three, per actor:
async def run(self) -> None:
while True:
enqueued, msg = await self.mailbox.get()
MAILBOX_WAIT.observe(time.perf_counter() - enqueued)
MAILBOX_DEPTH.set(self.mailbox.qsize())
await self.handle(msg)
Depth pinned at capacity means sustained overload. Wait near capacity × service time confirms it and tells you the latency callers are paying. A rising drop or rejection count tells you how much work is being shed. When all three move together, the actor needs more capacity — faster handling, batching, or more actors through sharding state by key — not a bigger mailbox, which only raises the latency ceiling.
Verify: a load test at twice capacity triggers the depth and drop alerts within one evaluation interval.
Verification¶
An actor's mailbox is safe under load when:
- The mailbox has a capacity chosen as tolerable latency divided by service time.
- The overload policy matches the message type: block for must-process work, reject for refusable requests, drop oldest or coalesce for state.
- Blocking
putcalls have a timeout, so senders cannot pile up indefinitely. - Depth, wait and drops are exported and alerted on together.
Diagnostic Hook: compare mailbox wait p99 with capacity × median service time. If wait is well below that, the actor has headroom; if it sits at that value, the actor is saturated and the mailbox is doing its job; if it is far above it, the mailbox is not actually bounded, or service time has a heavy tail worth investigating.
Pitfalls & edge cases¶
- Unbounded mailboxes. Measured: 4,338 queued and 4.6 s p99 wait at 2x load.
- Huge capacities. They behave like unbounded ones with extra steps; latency rises to match.
- Blocking
putwithout a timeout. Overload moves upstream and then has nowhere to stop. - Dropping silently. Count every drop, or overload looks like success.
Frequently Asked Questions¶
Should an asyncio actor's mailbox be bounded?
Yes. At twice the actor's capacity, an unbounded mailbox grew to 4,338 messages with a 4.6 s p99 wait; a 200-slot mailbox held p99 wait near 0.2 s under every overload policy.
What should happen when an actor's mailbox is full?
Block the sender (with a timeout) when every message must be processed, reject the newest when callers can be told no, and drop the oldest or coalesce by key when newer messages supersede older ones.
How big should an actor's mailbox be?
Capacity times service time is the worst-case wait, so divide the latency you can tolerate by the per-message handling time: 200 slots at about 1 ms per message gave the measured 220 ms p99 wait.
What is a coalescing mailbox?
A mailbox that keeps only the latest message per key, using a dict of latest values plus a queue of keys; its size is bounded by the number of distinct keys and the actor always sees the newest version.
Related¶
- Actors & Supervision — up to the topic overview.
- Sharding state across actors by key — adding capacity instead of queueing.
- Concurrent Execution & Worker Patterns — the section overview.