Resuming Interrupted Streaming Downloads¶
Retrying a large download from the beginning repeats everything already received; on a flaky link, a big enough file may never finish. HTTP range requests let a retry ask only for the missing bytes. Tested with httpx against a local server that cut the connection after 15, 12 and 10 MiB of a 40 MiB file: restarting from zero on each failure took four attempts and transferred 77 MiB; resuming with a Range header transferred exactly 40 MiB in the same four attempts. Resuming has one trap. When the server ignored the Range header and sent the whole file with status 200, a client that checked for 206 started over and still produced an intact file; a client that appended whatever arrived produced a 77 MiB file that was corrupt. This guide resumes safely: request the remainder, verify the response is the part you asked for and from the same version, and fall back to a full download when it is not.
Prerequisites¶
- Python 3.11+,
pip install httpx. - Streaming responses, from streaming large responses with httpx.
- Retry policy, from exponential backoff with jitter in asyncio.
1. Stream to a partial file and keep it on failure¶
The first requirement is to keep what was received. Stream to a .part file and do not delete it when the connection fails:
import os
import httpx
async def download_once(client: httpx.AsyncClient, url: str, part: str, headers: dict) -> httpx.Response:
async with client.stream("GET", url, headers=headers) as response:
response.raise_for_status()
mode = "ab" if response.status_code == 206 else "wb" # see step 3
with open(part, mode) as f:
async for chunk in response.aiter_bytes(1 << 20):
f.write(chunk)
return response
A failure mid-stream raises (httpx.ReadError, RemoteProtocolError, a timeout) after some chunks are on disk, and the partial file's size is exactly the offset to resume from. Write in large chunks; for very fast links, move the writes to a thread, as in streaming object downloads to disk without buffering.
Verify: after an interrupted attempt, the .part file's size equals the bytes received before the failure.
2. Ask for the remainder with Range and If-Range¶
On a retry, request from the partial file's size onwards, and make the request conditional on the resource being unchanged:
async def download(client: httpx.AsyncClient, url: str, dest: str, attempts: int = 6) -> None:
part = dest + ".part"
validator: str | None = None # ETag or Last-Modified from the first response
for attempt in range(attempts):
have = os.path.getsize(part) if os.path.exists(part) else 0
headers = {}
if have and validator:
headers = {"Range": f"bytes={have}-", "If-Range": validator}
try:
response = await download_once(client, url, part, headers)
validator = response.headers.get("ETag") or response.headers.get("Last-Modified")
break
except httpx.TransportError:
await asyncio.sleep(min(0.5 * 2 ** attempt, 10))
else:
raise RuntimeError(f"download of {url} failed after {attempts} attempts")
os.replace(part, dest)
Measured: three interruptions, four attempts, 40 MiB transferred for a 40 MiB file — no byte sent twice — against 77 MiB when every attempt restarted from zero. If-Range tells the server: send the range only if the resource still has this ETag, otherwise send the whole new version. That prevents stitching the first half of version 1 onto the second half of version 2. Capture the validator from the very first response (even a failed one, once headers have arrived) so the first retry can use it.
Verify: interrupt a download several times in a test; the total bytes transferred equal the file size.
3. Check the status before appending¶
A server may answer a range request with the full content — because it does not support ranges, because If-Range no longer matches, or because a proxy stripped the header. The status code says which:
def write_mode(response: httpx.Response, have: int) -> str:
if response.status_code == 206:
start = int(response.headers["Content-Range"].split()[1].split("-")[0])
if start != have:
raise ValueError(f"server resumed at {start}, expected {have}")
return "ab" # append the requested range
if response.status_code == 200:
return "wb" # full content: start the file over
raise ValueError(f"unexpected status {response.status_code}")
Tested: with a server that ignored Range, a client that switched to "wb" on 200 restarted correctly and ended with an intact file (transferring 77 MiB, as without resume); a client that appended regardless wrote the full file after the partial one and ended with a 77 MiB corrupt file. Checking Content-Range against the expected offset catches the rarer case of a server resuming at a different position. A 416 (range not satisfiable) usually means the partial file is already complete or larger than the resource — verify its size against the resource's length and start over if in doubt.
Verify: a test server that ignores Range still yields an intact file.
4. Verify the result¶
Resuming stitches several responses together, so verify the assembled file before using it:
import hashlib
def verify(path: str, expected_size: int | None, expected_sha256: str | None) -> None:
size = os.path.getsize(path)
if expected_size is not None and size != expected_size:
raise ValueError(f"size {size} != {expected_size}")
if expected_sha256:
digest = hashlib.sha256()
with open(path, "rb") as f:
for block in iter(lambda: f.read(1 << 20), b""):
digest.update(block)
if digest.hexdigest() != expected_sha256:
raise ValueError("checksum mismatch")
Take the expected size from the first response's Content-Length (or the total in Content-Range), and a checksum from wherever the publisher provides one — a checksum file, a header, the package index. Rename the .part file to its final name only after verification, so nothing downstream ever sees a stitched but wrong file.
Verify: a deliberately corrupted partial file is detected and the download restarts from zero.
5. Resume across process restarts¶
For very large files, keep the resume information on disk so a crashed or redeployed process can continue:
import json
def save_state(dest: str, url: str, validator: str | None, total: int | None) -> None:
with open(dest + ".part.json", "w") as f:
json.dump({"url": url, "validator": validator, "total": total}, f)
def load_state(dest: str) -> dict | None:
try:
with open(dest + ".part.json") as f:
return json.load(f)
except FileNotFoundError:
return None
With the URL, validator and total size stored next to the .part file, a new process can issue the same conditional range request. Discard the state if the URL changed. This is the same idea as checkpointing in checkpointing progress in long-running async jobs: the partial file is the checkpoint, and the validator guarantees it still belongs to the same resource.
Verify: kill the downloading process halfway; a new process completes the download transferring only the remaining bytes.
Verification¶
Resumable downloads are correct when:
- Partial data is kept in a
.partfile across failed attempts. - Retries send
RangeandIf-Rangewith the first response's validator. - Only 206 responses are appended, at the expected offset; 200 restarts the file.
- The result is verified before it gets its final name.
Diagnostic Hook: log, per download, attempts, bytes transferred and the file size. Bytes transferred much larger than the size means retries are restarting — the server may not support ranges, or the validator is missing; any verification failure after a resume points at a server or proxy that answers range requests incorrectly.
Pitfalls & edge cases¶
- Restarting from zero on large files. Measured: 77 MiB for 40 MiB.
- Appending without checking for 206. Tested: a corrupt 77 MiB file.
- Resuming without
If-Range. Two versions of the resource stitched together. - Renaming before verifying. A wrong file reaches downstream code.
Frequently Asked Questions¶
How do I resume an interrupted download with httpx?
Stream to a partial file, and on retry send Range: bytes=
What if the server ignores the Range header?
It answers 200 with the full content. Overwrite the partial file instead of appending; appending produced a corrupt 77 MiB file in testing.
What is If-Range for?
It makes the range request conditional: the server sends the range only if the resource still matches the given ETag or date, and otherwise sends the whole new version, preventing mixed versions.
How do I know a resumed download is correct?
Compare its size with the resource's total length and, when available, verify a published checksum before renaming the partial file to its final name.
Related¶
- Retry & Backoff Strategies — up to the topic overview.
- Retrying async calls with tenacity — the retry loop around each attempt.
- Resilience, Cancellation & Error Handling — the section overview.