Async File Access with anyio.Path¶
anyio.Path mirrors pathlib.Path with awaitable methods — read_text, write_bytes, exists, glob — that work on asyncio and trio. Each call runs the blocking operation in a worker thread, which keeps the event loop responsive but costs a thread round trip per call, and makes the call uncancellable while it is in the thread. Measured with AnyIO 4.15 on Python 3.14, reading 5,000 small JSON files of about 520 bytes from a RAM-backed filesystem: plain open().read() in a coroutine took 48–55 ms and blocked the loop for all of it. anyio.Path.read_text one file at a time took 317 ms on asyncio and 411 ms on trio with loop lag around 1 ms; aiofiles took 531 ms. Running all 5,000 reads concurrently was slower — 755–1,033 ms — with loop lag of 128–214 ms. A single to_thread.run_sync call that read every file took 50 ms with 0.9–1.1 ms of lag. And a read blocked on a FIFO with no writer ignored a 0.5 s move_on_after on both backends and was still blocked 5 s later; to_thread.run_sync(..., abandon_on_cancel=True) returned at 0.50 s. This guide shows where anyio.Path fits and where a coarser thread call is better.
Prerequisites¶
- AnyIO 4 on asyncio or trio.
- Worker threads in AnyIO, from calling sync code from AnyIO with to_thread and from_thread.
- The topic overview, AnyIO & Trio Interop.
1. Use anyio.Path for occasional file operations¶
For the file access a service does now and then — loading a template, writing a report, checking that a directory exists — anyio.Path is a drop-in, backend-neutral way to keep blocking calls off the loop:
from anyio import Path
async def load_config(path: str) -> dict:
p = Path(path)
if not await p.exists():
return {}
return json.loads(await p.read_text(encoding="utf-8"))
async def save_report(directory: str, name: str, body: bytes) -> None:
out = Path(directory) / name
await out.parent.mkdir(parents=True, exist_ok=True)
await out.write_bytes(body)
Path arithmetic — /, .name, .suffix, .parent — is synchronous, because it does no I/O; only methods that touch the filesystem are awaitable. Measured: 5,000 sequential read_text calls took 317 ms on asyncio and 411 ms on trio, about 63–82 µs per file, with loop lag around 1 ms; the same reads done synchronously on the loop took 48–55 ms but blocked it for that whole time. aiofiles, which also delegates to threads, took 531 ms. For a handful of calls per request the thread cost is noise; for thousands it dominates.
Verify: file access in request paths goes through anyio.Path or a thread call, and loop lag stays low during it.
2. Do not fan out thousands of tiny file operations¶
Running file reads concurrently looks like the async thing to do, but every call competes for AnyIO's default thread limiter of 40 tokens and for the GIL:
async with anyio.create_task_group() as tg:
for name in names: # 5,000 tasks, 40 threads, one GIL
tg.start_soon(read_one, name)
Measured: 755 ms on asyncio and 1,033 ms on trio for all 5,000 files at once — slower than reading them one at a time — with loop lag of 128 ms and 214 ms from scheduling thousands of tasks and thread completions. Bounding concurrency to 50 with a CapacityLimiter did not help: 811 ms and 1,099 ms. On a fast local filesystem, the work per file is tiny and the overhead per call is everything. Concurrency helps when each operation waits on a slow device or network filesystem; measure on the storage you actually use before parallelising.
There is a second trap in concurrent code like this: an augmented assignment whose right-hand side awaits.
total[0] += len(await Path(name).read_text()) # reads total[0] *before* the await
Measured: with 5,000 concurrent tasks, this summed to 522 bytes instead of 2,608,890, because each task loaded the old total, awaited, and wrote back its own sum. Await into a local first, then add.
Verify: concurrent file code is faster than the sequential version on your storage, and shared totals are updated after the await.
3. Batch file work into one thread call¶
When a task needs many file operations in a row, write them as ordinary synchronous code and run the whole function in one worker thread:
import anyio
def read_all(paths: list[str]) -> dict[str, str]:
out = {}
for p in paths:
with open(p, encoding="utf-8") as fh:
out[p] = fh.read()
return out
async def load_bundle(paths: list[str]) -> dict[str, str]:
return await anyio.to_thread.run_sync(read_all, paths)
Measured: 50 ms for all 5,000 files on both backends, with the loop's lag at 0.9–1.1 ms — the speed of plain synchronous code and the responsiveness of anyio.Path, because there is one thread hop instead of 5,000. The same applies to directory walks: anyio.Path.glob over the 5,000 files took 228 ms on asyncio and 341 ms on trio because it hops per entry, while os.scandir inside one to_thread call does the walk at native speed. Keep the function free of event-loop objects; if it must report progress back to async code, use anyio.from_thread.run sparingly.
Verify: bulk file operations run as one thread call, and their duration matches the synchronous equivalent.
4. Know that file calls cannot be interrupted¶
anyio.Path methods run in a worker thread that cannot be stopped from outside, and by default AnyIO waits for the thread to finish even when the calling task is cancelled. A file call that never returns — a hung NFS mount, a FIFO with no writer, a FUSE filesystem that has stopped responding — takes its cancel scope with it:
with anyio.move_on_after(0.5):
await Path("/mnt/stale-nfs/report.csv").read_text() # deadline cannot fire
Measured with a FIFO that nobody wrote to: on asyncio and on trio, the 0.5 s deadline did not end the call, and the process was still blocked when the test was killed after 5 s. When an operation might hang, call it through to_thread.run_sync with abandon_on_cancel=True, which lets the task stop waiting:
text = await anyio.to_thread.run_sync(read_report, path, abandon_on_cancel=True)
Measured: the deadline fired at 0.50 s on both backends. The thread itself is still blocked — abandoning only stops waiting for it — and it still holds a slot of the thread limiter, so a mount that keeps hanging will eventually exhaust all 40 slots and stall every anyio.Path call in the process. Treat abandoned threads as a symptom and alert on them, as in cancelling asyncio.to_thread calls.
Verify: file reads on storage that can hang use abandon_on_cancel=True under a deadline, and abandoned calls are counted.
5. Size the thread limiter for your file workload¶
All of AnyIO's thread calls — anyio.Path, open_file, to_thread.run_sync without an explicit limiter — share one default CapacityLimiter of 40 tokens per event loop. File access competes with every other blocking call routed through it:
limiter = anyio.to_thread.current_default_thread_limiter()
limiter.total_tokens = 80 # more slots for slow storage
FILE_LIMIT = anyio.CapacityLimiter(8) # or a separate, smaller pool for file I/O
async def read_slow(path: str) -> bytes:
return await anyio.to_thread.run_sync(read_bytes, path, limiter=FILE_LIMIT, abandon_on_cancel=True)
Measured here: raising concurrency did not speed up reads from fast storage, so the default is rarely too small for local disks. Give slow or unreliable storage its own small limiter, so hung or slow file calls cannot occupy the slots that DNS lookups, password hashing and other thread work depend on. For writing files safely — temporary file, flush, rename — the same thread-call pattern applies, as described in writing files atomically from async code.
Verify: slow or remote storage uses its own limiter, and the default limiter's borrowed tokens stay below its total under load.
Verification¶
File access from AnyIO code is efficient and safe when:
- Single operations use
anyio.Path, and none run synchronously on the loop. - Bulk operations run as one thread call, not thousands of awaitable ones.
- Reads from storage that can hang use
abandon_on_cancel=Trueunder a deadline. - Slow storage has its own
CapacityLimiter, separate from the default 40.
Diagnostic Hook: sample current_default_thread_limiter().borrowed_tokens and alert when it stays at the total. A limiter pinned at 40 with low CPU use usually means threads blocked in file or network calls that will not return — often a storage problem that surfaces first as every unrelated thread call in the service slowing down.
Pitfalls & edge cases¶
- Fanning out tiny file reads. Measured: 755–1,033 ms against 317–411 ms sequentially.
x += len(await ...)across tasks. Measured: 522 bytes summed instead of 2,608,890.- Deadlines around
anyio.Pathcalls. Measured: a blocked read ignored a 0.5 s deadline. - Abandoned threads. They keep their limiter slot until the call returns.
Frequently Asked Questions¶
Is anyio.Path really asynchronous?
It runs each filesystem call in a worker thread, so the loop stays responsive but each call costs a thread round trip: about 63 to 82 µs per small file read in testing.
Is it faster to read many files concurrently with anyio.Path?
Not on fast local storage: 5,000 concurrent reads took 755 to 1,033 ms against 317 to 411 ms one at a time. One to_thread call reading them all took 50 ms.
Can I cancel an anyio.Path read?
Not while it runs: a read blocked on a FIFO ignored a 0.5 s deadline. Use to_thread.run_sync with abandon_on_cancel=True, which returned at 0.50 s but leaves the thread blocked.
Should I use anyio.Path or aiofiles?
In AnyIO code, anyio.Path: it works on asyncio and trio and was faster here (317 ms against 531 ms for 5,000 files).
Related¶
- AnyIO & Trio Interop — up to the topic overview.
- Running subprocesses with AnyIO — the other blocking resource AnyIO wraps.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.