Skip to content

Batching Individual Calls with Futures

Code that loads one item at a time — await users.get(id) inside a loop, a resolver per field, a template that fetches each related record — produces a call per item, even when a backend could serve them all in one request. A batcher hides that: each caller gets a future for its key, the batcher collects the keys that arrive in the same moment, makes one batch call, and resolves every future from its result. Measured on Python 3.14 with a backend that took 10 ms per call (plus 0.1 ms per key in a batch) and 1,000 concurrent lookups over 600 distinct keys: individual calls, 50 at a time, made 1,000 backend calls, took 214 ms and had a p99 latency of 209 ms. A batcher that dispatched on the next loop tick with batches of at most 100 made 10 calls, took 32 ms and had a p99 of 27.5 ms. With requests arriving about one per millisecond instead of in a burst, a same-tick batcher made one call per request; a 2 ms window halved the calls and a 5 ms window cut them to a fifth, at 1.8–5.2 ms more latency per call. Without asyncio.shield, one cancelled caller cancelled the shared future and both callers of that key failed. This guide builds the batcher.

Prerequisites

1. Give each caller a future and schedule one dispatch

The batcher's load(key) does no I/O itself. It creates (or reuses) a future for the key, makes sure a dispatch is scheduled, and returns the future for the caller to await:

class Batcher:
    def __init__(self, batch_fn, max_batch: int = 100, window: float = 0.0):
        self.batch_fn, self.max_batch, self.window = batch_fn, max_batch, window
        self.pending: dict = {}                       # key -> future; repeated keys share one
        self.scheduled = False

    def load(self, key):
        loop = asyncio.get_running_loop()
        fut = self.pending.get(key)
        if fut is None:
            fut = loop.create_future()
            self.pending[key] = fut
            if len(self.pending) >= self.max_batch:
                self._dispatch()                       # full: send now
            elif not self.scheduled:
                self.scheduled = True
                if self.window:
                    loop.call_later(self.window, self._dispatch)
                else:
                    loop.call_soon(self._dispatch)     # after everything runnable this tick
        return asyncio.shield(fut)

call_soon runs the dispatch after every task already runnable in this loop iteration has had its turn, so all the load calls made by a burst of concurrent tasks land in the same batch. Keying the pending map by key also deduplicates: 1,000 lookups over 600 distinct keys produced batches totalling 918 keys, because keys requested twice in the same batch shared a future.

Verify: a burst of concurrent load calls produces one batch call per max_batch keys.

2. Resolve every future from the batch result

The dispatch takes the pending map, starts one task for the batch call, and resolves each future with its own result or error:

    def _dispatch(self):
        self.scheduled = False
        if not self.pending:
            return
        batch, self.pending = self.pending, {}
        asyncio.get_running_loop().create_task(self._run(batch))

    async def _run(self, batch: dict):
        try:
            results = await self.batch_fn(list(batch))
        except Exception as exc:
            for fut in batch.values():
                if not fut.done():
                    fut.set_exception(exc)              # every caller sees the batch failure
            return
        for key, fut in batch.items():
            if fut.done():
                continue
            if key in results:
                fut.set_result(results[key])
            else:
                fut.set_exception(KeyError(key))        # this key only

Measured: a missing key raised KeyError: 13 for its caller only, while other keys in the same batch resolved normally. Checking fut.done() before setting a result matters: a future may already have been cancelled, and setting a result on it raises InvalidStateError. Keep a reference to the batch task, or let a TaskGroup own it, in long-lived services — the loop holds tasks only weakly, as described in debugging unawaited coroutines in large codebases.

Verify: a failing batch call fails every caller in that batch, and a missing key fails only its own caller.

1,000 concurrent lookups over 600 keys A grid of 2 rows by 6 columns. 1,000 concurrent lookups over 600 keys approach backend calls mean batch total p50 p99 individual calls, 50 concurrent 1,000 1 214 ms 115.0 ms 208.7 ms batcher, same tick, max 100 10 91.8 32 ms 27.3 ms 27.5 ms Backend: 10 ms per call + 0.1 ms per key in a batch.

3. Choose a window for traffic that does not arrive in bursts

A same-tick batcher only batches calls that are made in the same loop iteration. Requests arriving one at a time from different clients rarely are:

batcher = Batcher(fetch_many, max_batch=100, window=0.002)    # wait up to 2 ms for company

Measured with 1,000 lookups arriving about one per millisecond: with no window, 1,000 backend calls and a p50 latency of 10.8 ms; with a 2 ms window, 500 calls and 12.6 ms; with 5 ms, 200 calls and 14.0 ms (p99 16.4 ms). The window is a direct trade: every call waits up to the window for others to join, in exchange for fewer, larger backend calls. Choose it from the backend's cost per call — a database round trip, a rate-limited API, a per-request fee — against the latency budget, and keep max_batch so that a burst still dispatches immediately when a batch fills.

Verify: under realistic arrival rates, backend calls per request and added latency are both measured for the chosen window.

Backend calls for 1,000 lookups arriving ~1 per ms 3 horizontal bars comparing window 0 (same tick), p50 10.8 ms with the others. Backend calls for 1,000 lookups arriving ~1 per ms window 0 (same tick), p50 10.8 ms 1,000 calls window 2 ms, p50 12.6 ms 500 calls window 5 ms, p50 14.0 ms 200 calls Arrivals spaced by asyncio.sleep(0.0005), about 1 ms apart in practice. Each millisecond of window buys fewer calls with latency.

4. Shield the shared future from one caller's cancellation

Several callers may await the same future — the same key requested twice, or deduplicated by design. If load returned the future itself, cancelling one caller's task would cancel the future, and every other caller of that key would get CancelledError:

def load_unshielded(self, key):
    fut = self._future_for(key)
    return fut                       # wrong: a caller's cancellation cancels the shared future

def load(self, key):
    fut = self._future_for(key)
    return asyncio.shield(fut)       # right: each caller can cancel only its own wait

Measured with two tasks loading key 1 and one of them cancelled before dispatch: without the shield, both tasks raised CancelledError; with it, the cancelled task stopped waiting and the other received {'id': 1}, from a single batch call of one key. The shared future still resolves when the batch returns, even if every caller has gone, so the backend work is not wasted on a retry; if that matters, count live waiters and drop keys with none before dispatching. Shielding is covered in more depth in using asyncio.shield to protect critical sections.

Verify: cancelling one of two callers of the same key leaves the other's result intact.

5. Scope batchers to a request or a short lifetime

A batcher that lives forever also caches nothing — its pending map empties at each dispatch — but it does mix requests: one tenant's keys can share a batch with another's, and an error in the batch fails both. For data with access control, create a batcher per request, which is also what the DataLoader pattern does in GraphQL servers, as in batching GraphQL resolvers with DataLoader:

async def handle(request):
    users = Batcher(lambda ids: fetch_users(ids, tenant=request.tenant))
    orders = await fetch_orders(request.tenant)
    async with asyncio.TaskGroup() as tg:
        tasks = [tg.create_task(users.load(o["user_id"])) for o in orders]   # one batch call
    return [t.result() for t in tasks]

A per-request batcher makes batching invisible to the code that calls load, keeps tenants apart, and leaves nothing behind when the request ends. For calls that should also be shared across requests and time — the same key fetched repeatedly — combine it with a cache or the single-flight pattern from implementing the single-flight pattern for duplicate calls.

Verify: batchers holding access-controlled data are created per request or per tenant, never shared across them.

Should these calls be batched, and how? A decision on How do the calls arrive with 4 outcomes. Should these calls be batched, and how? How do the calls arrive? burst within one request same-tick batcher 1,000 calls to 10 steady trickle window of a few ms 2 ms: half the calls same key, several callers shield each caller's wait no shared cancel tenant data batcher per request no cross-tenant batches max_batch bounds the size; the window bounds the wait.

Verification

A future-based batcher is correct when:

  • Concurrent load calls share one batch call, deduplicated by key.
  • Each future is resolved individually, with batch failures and missing keys reported to the right callers.
  • Callers' cancellations are shielded from the shared future.
  • The window and batch size are measured choices, and batchers are scoped per request where data is access-controlled.

Diagnostic Hook: export the batch size of every dispatch as a histogram. Batches of size one in a service that should be batching mean calls are not concurrent — usually because the caller awaits each load in a loop instead of creating them together — and the batcher is adding overhead without saving any calls.

Pitfalls & edge cases

  • Awaiting load in a loop. Each call dispatches alone; create the calls together.
  • Returning the shared future unshielded. Measured: one cancellation failed both callers.
  • A window for burst traffic. It only adds latency when calls already share a tick.
  • Process-wide batchers for tenant data. One batch can mix tenants and failures.

Frequently Asked Questions

How do I batch many asyncio calls into one request?

Give each caller a future for its key, schedule one dispatch with loop.call_soon, and resolve every future from the batch result. 1,000 concurrent lookups went from 1,000 backend calls to 10 in testing.

What batching window should I use?

None for burst traffic within a request; a few milliseconds for steady arrivals. At about one request per millisecond, a 2 ms window halved backend calls and added about 1.8 ms of latency.

Why did cancelling one caller fail the others?

They shared a future, and cancelling a task cancels the future it awaits. Return asyncio.shield(future) from load so each caller cancels only its own wait.

Is this the DataLoader pattern?

Yes: DataLoader is a per-request batcher with deduplication and usually a per-request cache.