Uploading Large Files to S3 with Async Multipart¶
A single put_object sends a file as one request, and the SDK holds the whole body in memory while it signs and sends it. For large files that is slow and memory-hungry, and one network error means starting over. Multipart upload splits the file into parts uploaded independently — several at a time — and assembles them on the server. Measured with aioboto3 15.5 uploading a 256 MiB file to a local S3-compatible server: a single put_object ran at 105 MiB/s and peaked at 579 MiB of process memory; a manual multipart upload with 8 MiB parts peaked at 91 MiB with one part in flight (77 MiB/s) and 204 MiB with eight (131 MiB/s); the managed upload_fileobj with the same part size and concurrency ran at 205 MiB/s with a 156 MiB peak. An upload that failed partway stayed listed as incomplete until it was aborted. This guide uses the managed transfer, shows the manual protocol for when you need control, and cleans up failures.
Prerequisites¶
- Python 3.11+,
pip install aioboto3; measured with aioboto3 15.5 / aiobotocore 2.25. - A shared client, from managing aioboto3 clients without leaking connections.
- S3's multipart rules: parts of at least 5 MiB except the last, at most 10,000 parts per upload.
1. Use the managed transfer for most uploads¶
upload_fileobj and upload_file choose single or multipart upload by size, upload parts concurrently, retry failed parts and abort on failure:
from boto3.s3.transfer import TransferConfig
UPLOADS = TransferConfig(
multipart_threshold=16 * 1024 * 1024, # files above 16 MiB use multipart
multipart_chunksize=8 * 1024 * 1024, # part size
max_concurrency=8, # parts in flight
)
async def upload(s3, path: str, bucket: str, key: str) -> None:
with open(path, "rb") as f:
await s3.upload_fileobj(f, bucket, key, Config=UPLOADS,
ExtraArgs={"ContentType": "application/octet-stream"})
Measured: 205 MiB/s with a 156 MiB peak for a 256 MiB file — faster than the single put and at about a quarter of its memory. Memory is bounded by part size × concurrency plus overhead, independent of file size. Keep max_concurrency within the client's max_pool_connections, or parts queue for connections inside the pool.
Verify: uploading a file several times larger than available memory succeeds, and process memory stays near part size × concurrency.
2. Choose part size and concurrency¶
Part size trades request count against memory and retry cost; concurrency trades throughput against memory and connections:
import math
def part_size_for(file_size: int, minimum: int = 8 * 2**20) -> int:
"""Smallest part size >= minimum that keeps the upload under 10,000 parts."""
needed = math.ceil(file_size / 10_000)
size = max(minimum, needed)
return math.ceil(size / 2**20) * 2**20 # round up to whole MiB
part_size_for(50 * 2**30) # 50 GiB -> about 5.12 MiB minimum, so 8 MiB parts (6,400 parts)
part_size_for(200 * 2**30) # 200 GiB -> 21 MiB parts
Over a real network, each part pays a round trip, so more concurrency helps until bandwidth saturates; measured locally, going from one part in flight to eight raised throughput from 77 to 131 MiB/s for the manual uploader. Parts smaller than 5 MiB are rejected at completion — tested, 1 KiB parts failed with EntityTooSmall. A failed part costs one part's worth of re-upload, which is the argument against very large parts on unreliable links.
Verify: the computed part count for your largest expected file stays below 10,000.
3. Drive the protocol yourself when you need control¶
The managed transfer hides the upload id, which is what you need to resume across process restarts, to upload parts from different sources, or to compute per-part checksums. The manual protocol is three calls:
async def multipart_upload(s3, path: str, bucket: str, key: str,
part_size: int = 8 * 2**20, concurrency: int = 8) -> None:
upload = await s3.create_multipart_upload(Bucket=bucket, Key=key)
upload_id = upload["UploadId"]
slots = asyncio.Semaphore(concurrency)
size = os.path.getsize(path)
def read_part(offset: int) -> bytes:
with open(path, "rb") as f:
f.seek(offset)
return f.read(part_size)
async def send(number: int, offset: int) -> dict:
async with slots:
body = await asyncio.to_thread(read_part, offset) # disk read off the loop
response = await s3.upload_part(Bucket=bucket, Key=key, UploadId=upload_id,
PartNumber=number, Body=body)
return {"PartNumber": number, "ETag": response["ETag"]}
try:
parts = await asyncio.gather(*(send(i + 1, off) for i, off in enumerate(range(0, size, part_size))))
await s3.complete_multipart_upload(Bucket=bucket, Key=key, UploadId=upload_id,
MultipartUpload={"Parts": parts})
except BaseException:
await asyncio.shield(s3.abort_multipart_upload(Bucket=bucket, Key=key, UploadId=upload_id))
raise
Part numbers start at 1 and the completion list must be in ascending order, which gather preserves. Reading each part inside the semaphore keeps at most concurrency parts in memory. The except BaseException covers cancellation too — a cancelled upload is a failed upload — and the shield lets the abort finish even while the task is being cancelled.
Verify: cancelling the task mid-upload leaves no incomplete upload listed for the key.
4. Clean up incomplete uploads¶
Parts of an upload that was never completed or aborted remain stored — invisible as an object, but billed. Tested: after uploading two parts without completing, list_multipart_uploads listed the upload and head_object for the key failed; only abort_multipart_upload removed it. Abort in code, and expire leftovers by policy for the crashes code cannot handle:
async def abort_stale_uploads(s3, bucket: str, older_than: timedelta) -> int:
cutoff = datetime.now(timezone.utc) - older_than
aborted = 0
paginator = s3.get_paginator("list_multipart_uploads")
async for page in paginator.paginate(Bucket=bucket):
for upload in page.get("Uploads", []):
if upload["Initiated"] < cutoff:
await s3.abort_multipart_upload(Bucket=bucket, Key=upload["Key"],
UploadId=upload["UploadId"])
aborted += 1
return aborted
# Bucket lifecycle rule (set once): AbortIncompleteMultipartUpload after 7 days
A lifecycle rule with AbortIncompleteMultipartUpload is the backstop on AWS and most compatible stores: it removes uploads that every code path missed, including those of processes that were killed. Keep the threshold longer than your longest legitimate upload, or a slow upload will be aborted under its feet.
Verify: list_multipart_uploads on the bucket returns only uploads currently in progress.
5. Resume interrupted uploads¶
For very large files over unreliable links, keep the upload id and the completed parts durably, and continue after a restart instead of starting over:
async def resume(s3, bucket: str, key: str, upload_id: str, path: str, part_size: int):
done = {}
paginator = s3.get_paginator("list_parts")
async for page in paginator.paginate(Bucket=bucket, Key=key, UploadId=upload_id):
for part in page.get("Parts", []):
done[part["PartNumber"]] = part["ETag"]
size = os.path.getsize(path)
todo = [(i + 1, off) for i, off in enumerate(range(0, size, part_size)) if i + 1 not in done]
log.info("resuming %s: %d parts done, %d to go", key, len(done), len(todo))
... # upload `todo` as in step 3, then complete
list_parts asks the server which parts it already has, so the local record only needs the upload id. The part size must be the same as in the original run, since part numbers map to file offsets. Store the upload id next to the file's identity (path, size, modification time) so a changed file starts a new upload instead of mixing contents.
Verify: kill an upload halfway, restart it, and only the missing parts are sent; the completed object's checksum matches the file.
Verification¶
Large uploads are handled well when:
- Files above a threshold use multipart, with memory bounded by part size × concurrency.
- Part size keeps uploads under 10,000 parts and at least 5 MiB.
- Failures and cancellations abort the upload, and a lifecycle rule expires leftovers.
- Huge uploads can resume from
list_partswith a stored upload id.
Diagnostic Hook: export upload throughput, parts retried, and the number of incomplete uploads per bucket over time. A steadily growing count of incomplete uploads means an abort path is missing; many retried parts mean the part size is too large for the link's error rate.
Pitfalls & edge cases¶
put_objectfor large files. Measured: 579 MiB peak for a 256 MiB file.- Parts under 5 MiB. Tested: completion fails with
EntityTooSmall. - Concurrency above the pool size. Parts wait for connections inside the client.
- No abort and no lifecycle rule. Orphaned parts are stored and billed indefinitely.
Frequently Asked Questions¶
How do I upload a large file to S3 with aioboto3?
Use await s3.upload_fileobj(file, bucket, key, Config=TransferConfig(...)) with a part size such as 8 MiB and a concurrency of around 8. In testing that ran at 205 MiB/s with a 156 MiB memory peak for a 256 MiB file.
What part size should I use for S3 multipart upload?
At least 5 MiB (the minimum for every part but the last), large enough to stay under 10,000 parts for your biggest file, and small enough that re-sending a failed part is cheap. 8 to 64 MiB is typical.
What happens to a multipart upload that is never completed?
Its parts stay stored and billed but no object appears. Abort it with abort_multipart_upload, and add a lifecycle rule to expire incomplete uploads automatically.
Can a multipart upload be resumed after a crash?
Yes, if you kept the upload id: list_parts returns the parts already stored, so you upload only the missing ones with the same part size and then complete.
Related¶
- Cloud SDKs & Object Storage — up to the topic overview.
- Streaming object downloads to disk without buffering — the download direction.
- Network I/O & Protocol Handling — the section overview.