Using MongoDB with PyMongo's Async API¶
PyMongo 4.x ships an asyncio API, AsyncMongoClient, which replaces the separate Motor package for new code. The method names match the synchronous driver, so the code looks familiar; its behaviour under concurrency has three traps worth measuring. Measured on Python 3.14 with PyMongo 4.18.2 against MongoDB 8.2 in Docker (8.0 refused to start on a Linux 7.0 kernel), with a 100,000-document collection: iterating a cursor with default batching stalled the event loop for up to 410 ms, while batch_size=1000 kept the longest stall to 8.6 ms and finished faster, 0.21 s against 0.51 s. With 200 tasks sharing a 10-connection pool, the median query took 1.3 ms, but the slowest waited 4,020 ms — the whole run — and some tasks completed 1 query while others completed 2,857; an asyncio.Semaphore(10) in front of the client made it fair, with a maximum of 36 ms and every task completing 153–154. A client created per request cost 2.44 ms per query against 0.20 ms shared, and insert_many loaded 100,000 documents in 0.52 s, where insert_one in a loop would take about 20 s. This guide covers each.
Prerequisites¶
- PyMongo 4.13 or later (
pip install pymongo), which includesAsyncMongoClient; a MongoDB server. - Connection pool basics, from sizing async connection pools for throughput.
- The topic overview, Async Database Drivers.
1. Create one client for the application¶
AsyncMongoClient owns a connection pool and background monitoring tasks. Create it once, in the application's lifespan, and close it on shutdown:
from pymongo import AsyncMongoClient
@asynccontextmanager
async def lifespan(app):
client = AsyncMongoClient("mongodb://db:27017", maxPoolSize=50)
app.state.db = client.app
yield
await client.close()
Measured: creating a client, running one find_one and closing it took 2.44 ms per query, against 0.20 ms with a shared client — the per-request client paid for a TCP connection and the server handshake every time. The client must be created and used inside a running event loop, and closed with await client.close(); it cannot be shared between loops, so tests that call asyncio.run repeatedly need a client per test.
Verify: exactly one AsyncMongoClient is constructed per process, and close() is awaited at shutdown.
2. Write in bulk¶
Each insert_one is a round trip. For loading or migrating data, send batches:
await col.insert_many(docs, ordered=False) # 100,000 documents
await col.bulk_write([
UpdateOne({"_id": d["_id"]}, {"$set": {"total": d["total"]}}, upsert=True)
for d in changes
], ordered=False)
Measured: insert_one in a loop took 200 µs per document, which extrapolates to 20 s for 100,000; insert_many inserted all 100,000 in 0.52 s, and with ordered=False 0.51 s. The driver splits large batches into server-sized messages itself. ordered=False makes little difference to speed on a single server, but it means one duplicate key does not stop the remaining inserts; with the default ordered=True, the first failure ends the batch.
Verify: loaders use insert_many or bulk_write, and the choice of ordered matches what should happen after the first error.
3. Set batch_size on large cursors¶
A cursor fetches documents in batches and decodes each batch's BSON on the event loop. With default settings, the first batch is 101 documents and later batches are as large as the server allows — up to 16 MiB — so decoding one batch can take hundreds of milliseconds:
async for order in col.find({"status": "open"}, batch_size=1000):
await handle(order)
Measured over 100,000 documents, with a probe task timing 1 ms sleeps: default batching took 0.51–0.55 s and stalled the loop for up to 380–410 ms; to_list() stalled it for 294–331 ms; batch_size=1000 stalled it for at most 8.6 ms and finished in 0.21–0.22 s. Smaller batches meant more round trips but less memory churn per batch, and here the result was faster as well as smoother. For endpoints that return a page of results, combine limit() with a matching batch_size.
Verify: every cursor that can return more than a few thousand documents sets batch_size, and a probe shows loop stalls below 10 ms while it runs.
4. Keep the pool fair under high concurrency¶
When more tasks than maxPoolSize query at once, they wait inside the driver for a connection. Measured with 200 tasks and maxPoolSize=10, each running find_one on an indexed field in a loop for 4 seconds: throughput was 7,186 queries per second and the median 1.3 ms — but the p99.9 and maximum were 4,020 ms, and per-task counts ranged from 1 to 2,857. Some tasks got a connection repeatedly while others waited for the whole run. A larger pool did not remove it: with maxPoolSize=200, the maximum was still 3,658 ms after a warm-up. A first-in-first-out semaphore in front of the client, sized to the pool, fixes the ordering:
db_slots = asyncio.Semaphore(10) # equal to maxPoolSize
async def find_order(customer: str):
async with db_slots:
return await col.find_one({"customer": customer})
Measured: 7,697 queries per second, a median of 25.8 ms, a p99 of 33.7 ms and a maximum of 36 ms, with every task completing 153–154 queries. The median rose because waiting became visible and evenly shared; the maximum fell by a factor of 110. For a service with a latency objective, the tail is what counts.
Verify: under a load test with more tasks than connections, the maximum latency is within a small multiple of the p99, and per-task completion counts are even.
5. Compare with threads before migrating¶
For a service already using the synchronous MongoClient in threads, measure before switching. With the same indexed find_one workload, 32 threads on a 32-connection pool reached 6,511 queries per second at 221 µs of client CPU per query; 100 threads reached 5,635 at 255 µs. The async client with a 10-connection pool reached 7,186–7,867 at 128–135 µs.
def sync_worker(col, stop):
r = random.Random()
while time.monotonic() < stop:
col.find_one({"customer": f"c{r.randrange(5000)}"})
The async API used about 40% less CPU per query and did not need more connections to go faster — throughput was highest with the smallest pool tested. For mixed applications, keep one client per concurrency model; a synchronous MongoClient used from async code blocks the loop on every call, as in running blocking SDK calls with asyncio.to_thread.
Verify: the migration's benefit is measured on the real query mix — CPU per query and tail latency — not assumed.
Verification¶
The async MongoDB client is used well when:
- One
AsyncMongoClientis shared per process and closed at shutdown. - Bulk writes use
insert_manyorbulk_write. - Large cursors set
batch_size, keeping loop stalls under 10 ms. - A semaphore sized to the pool bounds concurrent queries, so no task waits for seconds.
Diagnostic Hook: when a few requests take seconds while the median is a millisecond, count completions per task or per request handler under load. Counts from 1 to thousands — as measured with 200 tasks on 10 connections — mean waiters are not served in order, and a FIFO semaphore in front of the client will cap the wait.
Pitfalls & edge cases¶
- Default cursor batching on large results. Measured: 410 ms loop stalls.
- More tasks than connections with no semaphore. Measured: a 4,020 ms maximum.
- A client per request. Measured: 2.44 ms against 0.20 ms per query.
- MongoDB 8.0 on Linux kernels 6.19 and later. It refused to start here; 8.2 ran.
Frequently Asked Questions¶
Should I use Motor or PyMongo's async API?
PyMongo's AsyncMongoClient, for new code. It is part of the main driver from PyMongo 4.13, and its methods match the synchronous MongoClient.
Why does iterating a MongoDB cursor block my event loop?
Default batches can be up to 16 MiB, and each batch is decoded on the loop. Default batching stalled it for 410 ms; batch_size=1000 kept stalls under 8.6 ms.
Why do some MongoDB queries take seconds under load?
With more tasks than pool connections, some waited for the whole 4-second run. A FIFO asyncio.Semaphore sized to the pool cut the maximum to 36 ms.
Is async PyMongo faster than threads?
In this test, it used 128-135 µs of CPU per query against 221 µs for 32 threads, and reached 7,186-7,867 queries/s against 6,511.
Related¶
- Async Database Drivers — up to the topic overview.
- Using MySQL from asyncio — the same questions for MySQL drivers.
- Network I/O & Protocol Handling — the section overview.