Skip to content

Resuming Interrupted Streaming Downloads

Retrying a large download from the beginning repeats everything already received; on a flaky link, a big enough file may never finish. HTTP range requests let a retry ask only for the missing bytes. Tested with httpx against a local server that cut the connection after 15, 12 and 10 MiB of a 40 MiB file: restarting from zero on each failure took four attempts and transferred 77 MiB; resuming with a Range header transferred exactly 40 MiB in the same four attempts. Resuming has one trap. When the server ignored the Range header and sent the whole file with status 200, a client that checked for 206 started over and still produced an intact file; a client that appended whatever arrived produced a 77 MiB file that was corrupt. This guide resumes safely: request the remainder, verify the response is the part you asked for and from the same version, and fall back to a full download when it is not.

Prerequisites

1. Stream to a partial file and keep it on failure

The first requirement is to keep what was received. Stream to a .part file and do not delete it when the connection fails:

import os

import httpx


async def download_once(client: httpx.AsyncClient, url: str, part: str, headers: dict) -> httpx.Response:
    async with client.stream("GET", url, headers=headers) as response:
        response.raise_for_status()
        mode = "ab" if response.status_code == 206 else "wb"     # see step 3
        with open(part, mode) as f:
            async for chunk in response.aiter_bytes(1 << 20):
                f.write(chunk)
        return response

A failure mid-stream raises (httpx.ReadError, RemoteProtocolError, a timeout) after some chunks are on disk, and the partial file's size is exactly the offset to resume from. Write in large chunks; for very fast links, move the writes to a thread, as in streaming object downloads to disk without buffering.

Verify: after an interrupted attempt, the .part file's size equals the bytes received before the failure.

2. Ask for the remainder with Range and If-Range

On a retry, request from the partial file's size onwards, and make the request conditional on the resource being unchanged:

async def download(client: httpx.AsyncClient, url: str, dest: str, attempts: int = 6) -> None:
    part = dest + ".part"
    validator: str | None = None                        # ETag or Last-Modified from the first response
    for attempt in range(attempts):
        have = os.path.getsize(part) if os.path.exists(part) else 0
        headers = {}
        if have and validator:
            headers = {"Range": f"bytes={have}-", "If-Range": validator}
        try:
            response = await download_once(client, url, part, headers)
            validator = response.headers.get("ETag") or response.headers.get("Last-Modified")
            break
        except httpx.TransportError:
            await asyncio.sleep(min(0.5 * 2 ** attempt, 10))
    else:
        raise RuntimeError(f"download of {url} failed after {attempts} attempts")
    os.replace(part, dest)

Measured: three interruptions, four attempts, 40 MiB transferred for a 40 MiB file — no byte sent twice — against 77 MiB when every attempt restarted from zero. If-Range tells the server: send the range only if the resource still has this ETag, otherwise send the whole new version. That prevents stitching the first half of version 1 onto the second half of version 2. Capture the validator from the very first response (even a failed one, once headers have arrived) so the first retry can use it.

Verify: interrupt a download several times in a test; the total bytes transferred equal the file size.

Bytes transferred to download a 40 MiB file cut 3 times 4 horizontal bars comparing restart from zero with the others. Bytes transferred to download a 40 MiB file cut 3 times restart from zero 77 MiB, intact Range + If-Range resume 40 MiB, intact server ignores Range, client checks 206 77 MiB, intact server ignores Range, client appends blindly 77 MiB file, corrupt httpx against a local aiohttp server that aborted the connection after 15, 12 and 10 MiB. Resume saves the repeat; checking for 206 saves the file.

3. Check the status before appending

A server may answer a range request with the full content — because it does not support ranges, because If-Range no longer matches, or because a proxy stripped the header. The status code says which:

def write_mode(response: httpx.Response, have: int) -> str:
    if response.status_code == 206:
        start = int(response.headers["Content-Range"].split()[1].split("-")[0])
        if start != have:
            raise ValueError(f"server resumed at {start}, expected {have}")
        return "ab"                                   # append the requested range
    if response.status_code == 200:
        return "wb"                                   # full content: start the file over
    raise ValueError(f"unexpected status {response.status_code}")

Tested: with a server that ignored Range, a client that switched to "wb" on 200 restarted correctly and ended with an intact file (transferring 77 MiB, as without resume); a client that appended regardless wrote the full file after the partial one and ended with a 77 MiB corrupt file. Checking Content-Range against the expected offset catches the rarer case of a server resuming at a different position. A 416 (range not satisfiable) usually means the partial file is already complete or larger than the resource — verify its size against the resource's length and start over if in doubt.

Verify: a test server that ignores Range still yields an intact file.

4. Verify the result

Resuming stitches several responses together, so verify the assembled file before using it:

import hashlib


def verify(path: str, expected_size: int | None, expected_sha256: str | None) -> None:
    size = os.path.getsize(path)
    if expected_size is not None and size != expected_size:
        raise ValueError(f"size {size} != {expected_size}")
    if expected_sha256:
        digest = hashlib.sha256()
        with open(path, "rb") as f:
            for block in iter(lambda: f.read(1 << 20), b""):
                digest.update(block)
        if digest.hexdigest() != expected_sha256:
            raise ValueError("checksum mismatch")

Take the expected size from the first response's Content-Length (or the total in Content-Range), and a checksum from wherever the publisher provides one — a checksum file, a header, the package index. Rename the .part file to its final name only after verification, so nothing downstream ever sees a stitched but wrong file.

Verify: a deliberately corrupted partial file is detected and the download restarts from zero.

A resumable download, end to end A flow of 5 stages. A resumable download, end to end stream to .part remember ETag failure retry with backoff Range: bytes=size- If-Range: ETag 206 append / 200 restart check Content-Range verify size + hash rename Every step either saves work or prevents a corrupt result.

5. Resume across process restarts

For very large files, keep the resume information on disk so a crashed or redeployed process can continue:

import json


def save_state(dest: str, url: str, validator: str | None, total: int | None) -> None:
    with open(dest + ".part.json", "w") as f:
        json.dump({"url": url, "validator": validator, "total": total}, f)


def load_state(dest: str) -> dict | None:
    try:
        with open(dest + ".part.json") as f:
            return json.load(f)
    except FileNotFoundError:
        return None

With the URL, validator and total size stored next to the .part file, a new process can issue the same conditional range request. Discard the state if the URL changed. This is the same idea as checkpointing in checkpointing progress in long-running async jobs: the partial file is the checkpoint, and the validator guarantees it still belongs to the same resource.

Verify: kill the downloading process halfway; a new process completes the download transferring only the remaining bytes.

How should this download be retried? A decision on What does the server support, and how big is the file with 4 outcomes. How should this download be retried? What does the server support, and how big is the file? small file restart from zero simplest Accept-Ranges: bytes Range + If-Range 40 vs 77 MiB no range support restart, check 200 never append blindly huge file, long job persist .part + state survives restarts Resume when the server allows it, and always verify what you append.

Verification

Resumable downloads are correct when:

  • Partial data is kept in a .part file across failed attempts.
  • Retries send Range and If-Range with the first response's validator.
  • Only 206 responses are appended, at the expected offset; 200 restarts the file.
  • The result is verified before it gets its final name.

Diagnostic Hook: log, per download, attempts, bytes transferred and the file size. Bytes transferred much larger than the size means retries are restarting — the server may not support ranges, or the validator is missing; any verification failure after a resume points at a server or proxy that answers range requests incorrectly.

Pitfalls & edge cases

  • Restarting from zero on large files. Measured: 77 MiB for 40 MiB.
  • Appending without checking for 206. Tested: a corrupt 77 MiB file.
  • Resuming without If-Range. Two versions of the resource stitched together.
  • Renaming before verifying. A wrong file reaches downstream code.

Frequently Asked Questions

How do I resume an interrupted download with httpx?

Stream to a partial file, and on retry send Range: bytes=- with If-Range set to the ETag from the first response. Append only if the status is 206. In testing that transferred 40 MiB for a 40 MiB file cut three times.

What if the server ignores the Range header?

It answers 200 with the full content. Overwrite the partial file instead of appending; appending produced a corrupt 77 MiB file in testing.

What is If-Range for?

It makes the range request conditional: the server sends the range only if the resource still matches the given ETag or date, and otherwise sends the whole new version, preventing mixed versions.

How do I know a resumed download is correct?

Compare its size with the resource's total length and, when available, verify a published checksum before renaming the partial file to its final name.