Skip to content

Streaming Object Downloads to Disk Without Buffering

await response["Body"].read() is the shortest way to download an object and the worst way to download a large one: the whole object is held in memory, and then written out as a second step. Streaming it in chunks keeps memory constant and overlaps receiving with writing. Measured with aioboto3 15.5 downloading a 256 MiB object from a local S3-compatible server: read() followed by a write peaked at 577 MiB of process memory (from a 67 MiB baseline) and ran at 303 MiB/s; iterating the body in 1 MiB chunks and writing each through a thread peaked at 72 MiB and ran at 491 MiB/s; the managed download_file ran at 501 MiB/s and peaked at 189 MiB. This guide streams downloads to disk, verifies them, makes them atomic and resumable, and keeps disk writes off the event loop.

Prerequisites

1. Iterate the body in chunks

The response body is an async stream. Iterate it and write each chunk as it arrives:

CHUNK = 1024 * 1024


async def download(s3, bucket: str, key: str, path: str) -> int:
    response = await s3.get_object(Bucket=bucket, Key=key)
    written = 0
    with open(path, "wb") as f:
        async for chunk in response["Body"].iter_chunks(CHUNK):
            await asyncio.to_thread(f.write, chunk)          # disk write off the loop
            written += len(chunk)
    return written

Measured: 72 MiB peak (5 MiB above baseline) against 577 MiB for read(), and faster — 491 against 303 MiB/s — because writing overlaps receiving instead of following it. The iteration consumes the body to the end, which also releases the connection to the pool. With several downloads in parallel, memory is about chunk size × concurrency, independent of object sizes.

Verify: process memory during a download of an object larger than available RAM stays near its baseline.

Peak memory downloading a 256 MiB object 3 horizontal bars comparing body.read() then write with the others. Peak memory downloading a 256 MiB object body.read() then write 577 MiB, 303 MiB/s iter_chunks(1 MiB) to disk 72 MiB, 491 MiB/s managed download_file 189 MiB, 501 MiB/s aioboto3 15.5, local S3-compatible server; process baseline 67 MiB. Streaming is both leaner and faster than buffering the whole object.

2. Write to a temporary name and rename

A download interrupted halfway must never leave a truncated file under the final name, where the next step would treat it as complete:

import os


async def download_atomic(s3, bucket: str, key: str, path: str) -> int:
    tmp = f"{path}.part"
    try:
        written = await download(s3, bucket, key, tmp)
        await asyncio.to_thread(_fsync_path, tmp)
        os.replace(tmp, path)                               # atomic: whole file or nothing
        return written
    except BaseException:
        if os.path.exists(tmp):
            os.unlink(tmp)
        raise


def _fsync_path(path: str) -> None:
    fd = os.open(path, os.O_RDONLY)
    try:
        os.fsync(fd)
    finally:
        os.close(fd)

The .part file is invisible to anything that looks for the final name, and os.replace swaps it in only after the last byte is on disk. Catching BaseException removes the partial file on cancellation too. Keep the temporary file in the same directory as the target, so the rename stays on one filesystem.

Verify: cancel downloads at random points in a test; no file exists under a final name unless it is complete.

3. Verify what you received

Network transfers are checked by TCP, but truncation and application bugs are not. Compare the size, and a checksum when one is available:

import hashlib


async def download_verified(s3, bucket: str, key: str, path: str) -> None:
    response = await s3.get_object(Bucket=bucket, Key=key, ChecksumMode="ENABLED")
    expected_size = response["ContentLength"]
    expected_sha = response.get("ChecksumSHA256")             # present if uploaded with one
    digest = hashlib.sha256()
    tmp = f"{path}.part"
    with open(tmp, "wb") as f:
        async for chunk in response["Body"].iter_chunks(CHUNK):
            digest.update(chunk)
            await asyncio.to_thread(f.write, chunk)
    size = os.path.getsize(tmp)
    if size != expected_size:
        os.unlink(tmp)
        raise IOError(f"{key}: got {size} bytes, expected {expected_size}")
    if expected_sha and base64.b64encode(digest.digest()).decode() != expected_sha:
        os.unlink(tmp)
        raise IOError(f"{key}: checksum mismatch")
    os.replace(tmp, path)

Hashing in the loop costs CPU on the event loop — SHA-256 runs at roughly a gigabyte per second per core, so 1 MiB chunks take about a millisecond each; move hashing into the same thread call as the write if that matters. Multipart-uploaded objects have composite checksums or ETags that are not a plain hash of the content; store your own checksum in object metadata at upload time when you need end-to-end verification.

Verify: a test that truncates the body (through a fault-injecting proxy) raises the size error and leaves no final file.

A safe streaming download A flow of 5 stages. A safe streaming download get_object headers: size, checksum iter_chunks(1 MiB) hash + write in thread fsync .part durable verify size, checksum or delete os.replace final name Memory stays at one chunk; the final name only ever holds verified files.

4. Resume large downloads with Range requests

For multi-gigabyte objects over unreliable links, restarting from zero after a failure wastes everything already received. Ask for the remaining bytes only:

async def download_resumable(s3, bucket: str, key: str, path: str) -> None:
    tmp = f"{path}.part"
    have = os.path.getsize(tmp) if os.path.exists(tmp) else 0
    head = await s3.head_object(Bucket=bucket, Key=key)
    total, etag = head["ContentLength"], head["ETag"]
    if have >= total:
        os.replace(tmp, path)
        return
    response = await s3.get_object(Bucket=bucket, Key=key, Range=f"bytes={have}-", IfMatch=etag)
    with open(tmp, "ab") as f:                                 # append to what we have
        async for chunk in response["Body"].iter_chunks(CHUNK):
            await asyncio.to_thread(f.write, chunk)
    if os.path.getsize(tmp) != total:
        raise IOError("incomplete download; retry to resume")
    os.replace(tmp, path)

IfMatch with the ETag makes the request fail if the object changed since the partial download began, which prevents stitching two versions together; on that failure, delete the partial file and start over. Range requests also allow downloading one large object in parallel ranges, which the managed download_file does internally — it measured 501 MiB/s against 491 for a single stream, a small gain locally that grows on high-latency links.

Verify: interrupt a download, rerun it, and only the remaining bytes are requested; the final file's checksum matches the object.

5. Choose between your stream and the managed transfer

download_file and download_fileobj handle ranges, concurrency and retries for you. Your own stream gives control over hashing, verification and where bytes go:

from boto3.s3.transfer import TransferConfig

# Managed: parallel ranged GETs, retries, writes to a file
await s3.download_file(bucket, key, path,
                       Config=TransferConfig(multipart_chunksize=8 * 2**20, max_concurrency=8))

# Your stream: when bytes go somewhere other than a local file
response = await s3.get_object(Bucket=bucket, Key=key)
async for chunk in response["Body"].iter_chunks(CHUNK):
    await upstream_writer.write(chunk)                   # e.g. to an HTTP response, a parser, a pipe

Measured, the managed transfer was the fastest (501 MiB/s) at a higher memory peak (189 MiB) from its concurrent ranges; the single stream used the least memory. Streaming your own way is the only option when the destination is not a file: proxying an object to an HTTP client as in streaming responses with Starlette and FastAPI, feeding a decompressor or parser, or writing into another store.

Verify: the chosen method's throughput and memory peak, measured on your largest objects, fit your budgets.

How should this object be downloaded? A decision on How big, and where does it go with 4 outcomes. How should this object be downloaded? How big, and where does it go? small (MiB) body.read() is fine inside async with large, to a local file download_file or iter_chunks .part + rename huge, unreliable link Range resume + IfMatch keep the .part to another destination iter_chunks to it constant memory Never let memory grow with object size.

Verification

Downloads are streamed correctly when:

  • Bodies are iterated in chunks, with writes off the event loop.
  • Files appear under their final name only when complete, via .part and os.replace.
  • Size and, where available, checksums are verified before the rename.
  • Large downloads resume with Range and IfMatch.

Diagnostic Hook: record bytes per second and peak memory per download, and count verification failures and resumes. Memory that scales with object size means a code path still calls read(); frequent resumes for one source point at an unstable network path or too-short timeouts.

Pitfalls & edge cases

  • await body.read() for large objects. Measured: 577 MiB peak for 256 MiB.
  • Writing chunks synchronously on the loop. Disk latency stalls every other task.
  • Final names for partial files. Truncated files look complete to the next step.
  • Resuming without IfMatch. A changed object produces a stitched, corrupt file.

Frequently Asked Questions

How do I download a large S3 object to disk with aioboto3?

Call get_object, iterate response["Body"].iter_chunks(1024 * 1024), and write each chunk to a temporary file through asyncio.to_thread, renaming it when complete. A 256 MiB object peaked at 72 MiB of memory that way in testing.

Is iter_chunks slower than reading the whole body?

No; in testing it was faster, 491 against 303 MiB/s, because receiving and writing overlap.

How do I resume an interrupted S3 download?

Keep the partial file, request Range bytes from its current size to the end with IfMatch set to the original ETag, and append the result.

Should I use download_file or stream it myself?

download_file is fastest for local files thanks to parallel ranges (501 MiB/s in testing); stream it yourself when you need custom verification or the bytes go somewhere other than a file.