Skip to content

Processing Large CSV and JSONL Files Asynchronously

Large input files meet asyncio in two places: reading and parsing them without stalling the event loop, and then doing asynchronous work per record without creating a million tasks at once. Measured on Python 3.14 with a 1,000,000-line JSONL file of 119 MB, held in the page cache: reading it whole and parsing it in a coroutine took 1.62 s but blocked the event loop for all 1,622 ms and peaked at 297 MiB. Iterating it line by line with aiofiles kept the loop free — 2.7 ms maximum lag — and took 31.8 s, twenty times slower, because every line was a round trip to a thread. Reading and parsing 1 MiB chunks in asyncio.to_thread took 1.42 s with 10 ms of lag and 36 MiB. Splitting the file into byte ranges across an eight-process pool took 0.27 s. For per-record I/O, gathering one coroutine per record took 11.2 s, peaked at 2,040 MiB and stalled the loop for 4.0 s; a bounded reader-and-workers pipeline did the same work in 4.85 s at 40 MiB with 7.6 ms of lag. This guide shows how to read big files from async code and which approach fits which workload.

Prerequisites

1. Do not read and parse a big file on the loop

The simplest code reads the file and parses it inside a coroutine. It works, and it stops everything else the process is doing:

async def total_amount(path):
    with open(path) as f:
        data = f.read()                                        # 119 MB, on the loop
    return sum(json.loads(line)["amount"] for line in data.splitlines())

Measured: 1.62 s, a maximum loop lag of 1,622 ms — the entire run — and a peak RSS of 297 MiB, because the whole file and its list of lines were in memory at once. In a service, every request, heartbeat and timer waits for those 1.6 s. Even a small file read synchronously is a blocking call; it is just a shorter one, and the file was already in the page cache here — on a cold disk or a network filesystem the stall is longer. Reading belongs off the loop, and big files must be streamed rather than loaded whole.

Verify: loop lag measured during a file import stays within your service's budget.

Summing one field over 1,000,000 JSONL lines A grid of 4 rows by 5 columns. Summing one field over 1,000,000 JSONL lines approach time lines/s max loop lag peak RSS read + parse on the loop 1.62 s 617k 1,622 ms 297 MiB aiofiles, async for line 31.82 s 31k 2.7 ms 23 MiB 1 MiB chunks in to_thread 1.42 s 704k 10.0 ms 36 MiB 8 byte ranges, warm process pool 0.27 s 3.7M 5.4 ms 26 MiB* Python 3.14, file in the page cache. *Parent process only; each worker adds its own.

2. Read in chunks from a thread, not line by line

aiofiles makes file reads awaitable by running each one in a thread pool. Per line, that overhead dominates:

async with aiofiles.open(path) as f:
    async for line in f:                                       # one thread hop per line
        total += json.loads(line)["amount"]

Measured: 31.8 s, against 1.42 s for the same work in chunks. Move the hop to a coarser grain: read about a megabyte of whole lines, and parse them, in one to_thread call:

def read_chunk(f, size_hint=1 << 20) -> list[dict]:
    return [json.loads(line) for line in f.readlines(size_hint)]   # whole lines, ~1 MiB

async def iter_records(path):
    with open(path) as f:
        while records := await asyncio.to_thread(read_chunk, f):
            for record in records:
                yield record

readlines(hint) stops at a line boundary after about hint bytes, so chunks never split a record. Measured: 1.42 s, 704,000 lines/s, 36 MiB peak and a maximum loop lag of 10 ms. The lag comes from the coroutine's own work per chunk and from the parsing thread holding the GIL; smaller chunks lower it at some cost in throughput. The open file object is used by one thread at a time — each to_thread call finishes before the next starts — so sharing it is safe.

Verify: the reader yields every record exactly once, and memory stays flat as file size grows.

3. Read CSV rows from one reader, in batches

CSV adds a complication: a quoted field can contain newlines, so splitting the file on line boundaries can cut a record in half. Keep one csv reader over the file and pull a batch of rows from it per thread call:

import csv
import itertools

def next_rows(reader, n=10_000) -> list[dict]:
    return list(itertools.islice(reader, n))

async def iter_csv(path):
    with open(path, newline="") as f:
        reader = csv.DictReader(f)
        while rows := await asyncio.to_thread(next_rows, reader):
            for row in rows:
                yield row

Measured on the same million records as a 65 MB CSV file: 1.15 s, 868,000 rows/s, 32 MiB peak and 19.4 ms maximum loop lag. The reader's state — including a half-read quoted field — lives in the reader object between calls, so multi-line fields are handled correctly. Values arrive as strings; convert them in next_rows, inside the thread, so the conversion cost stays off the loop.

Verify: a CSV file with quoted newlines in a field produces the same rows as csv.DictReader read synchronously.

4. Use a process pool for parse-heavy files

When parsing dominates — JSON with large nested records, validation, transformation — threads do not help, because parsing holds the GIL. Split the file into byte ranges and parse them in parallel processes; each worker skips forward to the first full line in its range:

def parse_range(path, start, end) -> float:
    total = 0.0
    with open(path, "rb") as f:
        if start:
            f.seek(start - 1)
            f.readline()                       # finish the line that straddles start
        while f.tell() < end and (line := f.readline()):
            total += json.loads(line)["amount"]
    return total

async def total_parallel(path, pool, n=8):
    size = os.path.getsize(path)
    step = size // n
    loop = asyncio.get_running_loop()
    parts = await asyncio.gather(*(
        loop.run_in_executor(pool, parse_range, path, i * step, size if i == n - 1 else (i + 1) * step)
        for i in range(n)
    ))
    return sum(parts)

Measured with eight processes: 0.27 s and 5.4 ms of lag with a pool created at start-up; 0.33 s and 58.1 ms of lag when the pool was created and shut down inside the coroutine with a with ProcessPoolExecutor() block — its exit calls shutdown(wait=True), which blocks the event loop while workers exit. Create the pool once for the life of the service. The result matched the single-threaded sum exactly, which is the test that no line was skipped or counted twice at a boundary.

Verify: the parallel result equals the sequential one on a file whose size is not a multiple of the range count.

How should this file be read? A decision on What does the file need with 4 outcomes. How should this file be read? What does the file need? large JSONL, light parsing 1 MiB chunks in to_thread 1.42 s, 36 MiB CSV, possibly multi-line fields row batches from one reader 1.15 s, 32 MiB CPU-heavy parsing byte ranges in a warm process pool 0.27 s line-by-line aiofiles avoid for big files 31.8 s The grain of each thread or process hop decides the speed.

5. Do per-record I/O with bounded concurrency

Often each record needs an asynchronous call — an API lookup, a database write. The tempting code collects the records and gathers one coroutine per record behind a semaphore:

records = [r async for r in iter_records(path)]
sem = asyncio.Semaphore(500)

async def handle(record):
    async with sem:
        await lookup(record)

await asyncio.gather(*(handle(r) for r in records))          # a million tasks at once

The semaphore limits how many calls are in flight, not how many tasks exist. Measured with a 1 ms call per record: 11.2 s, a peak RSS of 2,040 MiB and a maximum loop lag of 4.0 s — creating and scheduling a million tasks is itself a long blocking operation. A pipeline with a bounded queue between the reader and a fixed set of workers creates no more objects than it is using:

async def process_file(path, workers=500):
    q: asyncio.Queue = asyncio.Queue(2000)

    async def read():
        async for record in iter_records(path):
            await q.put(record)                 # waits when workers fall behind
        for _ in range(workers):
            await q.put(None)

    async def work():
        while (record := await q.get()) is not None:
            await lookup(record)

    async with asyncio.TaskGroup() as tg:
        tg.create_task(read())
        for _ in range(workers):
            tg.create_task(work())

Measured: 4.85 s, 40 MiB and 7.6 ms of loop lag for the same million records. The reader blocks on put when the queue is full, so the file is read only as fast as records are handled — backpressure all the way to the disk. Errors per record are handled inside work, as in handling per-item errors in async pipelines.

Verify: peak memory for a 10× larger file is the same as for the original.

One async call per record, 1,000,000 records A grid of 2 rows by 4 columns. One async call per record, 1,000,000 records approach time peak RSS max loop lag gather(handle(r) for r in records), Semaphore(500) 11.20 s 2,040 MiB 3,968 ms bounded queue (2,000) + 500 workers 4.85 s 40 MiB 7.6 ms Each record's call was a 1 ms sleep; concurrency 500 in both.

Verification

Large files are processed safely from asyncio when:

  • No file is read or parsed on the loop, measured by loop lag during an import.
  • Reads are chunked, by bytes for JSONL and by rows for CSV, one thread hop per chunk.
  • CPU-heavy parsing runs in a long-lived process pool over byte ranges.
  • Per-record async work runs through a bounded queue with a fixed number of workers.

Diagnostic Hook: log records read per second and peak RSS for every import job. Throughput far below what a chunked reader achieves points at a per-line hop such as aiofiles; memory that grows with file size points at records collected into a list or a task per record — the two mistakes that this guide measured at 20× slower and 50× larger.

Pitfalls & edge cases

  • async for line in aiofiles.open(...). Measured: 31.8 s for what took 1.4 s in chunks.
  • Splitting CSV on newlines. Quoted fields can span lines; batch rows from one reader.
  • with ProcessPoolExecutor() inside a coroutine. Its shutdown blocked the loop for 58 ms.
  • A task per record. Measured: 2,040 MiB and a 4 s stall for a million records.

Frequently Asked Questions

How do I read a large file in asyncio without blocking?

Read and parse it in chunks inside asyncio.to_thread, yielding records to async code. A 119 MB JSONL file took 1.42 s with 10 ms of loop lag this way, against a 1.6 s stall when read on the loop.

Is aiofiles fast for reading large files?

Not line by line: every line is a thread round trip, and a million lines took 31.8 s against 1.42 s for 1 MiB chunks in a thread. Use it, if at all, for coarse reads.

How do I process a large CSV file asynchronously?

Keep one csv.DictReader over the open file and pull batches of rows with itertools.islice inside asyncio.to_thread; this handles multi-line quoted fields and read a million rows in 1.15 s.

How do I call an API for every record in a huge file?

Feed records through a bounded queue to a fixed number of worker coroutines. A million records took 4.85 s and 40 MiB this way, against 11.2 s and 2 GiB with one gathered coroutine per record.