Skip to content

Handling Redirects and Timeouts When Crawling with aiohttp

A crawler fetches from servers it knows nothing about, so every request must be bounded in hops, time and size — and bounded in ways that hostile or broken servers cannot get around. aiohttp provides the controls; their defaults and exact semantics matter. Tested with aiohttp 3.14 against a local test site: a redirect was followed transparently with its 302 recorded in response.history; a URL that redirected to itself raised TooManyRedirects after 10 hops, the default limit; and a server that sent one byte every half-second was still being read after 12 s with a sock_read timeout of 2 s, because each byte arrived well within 2 s. A total timeout of 5 s stopped it at 5.3 s. This guide sets each limit for crawling, decides what to do with redirects across hosts, and caps body size.

Prerequisites

1. Bound redirects and record where you landed

aiohttp follows redirects by default, up to max_redirects=10, and records each hop:

async with session.get(url, max_redirects=5) as response:
    final_url = str(response.url)                         # where the content actually came from
    hops = [(h.status, str(h.url)) for h in response.history]

Tested: /redirect/5 returned the content of /p/5 with history == [302]; a self-redirecting URL raised aiohttp.TooManyRedirects after 10 hops, and after 3 with max_redirects=3. Five is plenty for real sites — chains longer than that are usually loops or tracking hops. Always use response.url, not the requested URL, as the base for resolving relative links in the page and as the URL you record as fetched; otherwise links on redirected pages resolve against the wrong path.

Verify: a redirect loop on a test host fails fast with TooManyRedirects, and links on a redirected page resolve against the final URL.

What each limit did against misbehaving test endpoints A grid of 5 rows by 3 columns. What each limit did against misbehaving test endpoints endpoint setting result 302 to /p/5 default followed, history [302] redirect to itself default (10) TooManyRedirects after 10 hops redirect to itself max_redirects=3 TooManyRedirects after 3 1 byte every 0.5 s sock_read=2 still reading after 12 s 1 byte every 0.5 s total=5 TimeoutError at 5.3 s aiohttp 3.14 against a local test site.

2. Decide what cross-host redirects mean

A redirect can move the request to a different host — http to https, example.com to www.example.com, or a different site entirely. That matters for politeness, robots rules and scope:

async def fetch(session, url: str, scheduler) -> FetchResult | None:
    async with session.get(url, allow_redirects=False) as response:
        if response.status in (301, 302, 303, 307, 308):
            location = response.headers.get("Location")
            if location:
                target = normalize(urljoin(url, location))
                scheduler.offer(target)                   # re-enters the frontier: dedup, robots, host queue
            return None
        return await read_page(response)

Handling redirects yourself with allow_redirects=False sends each target back through the frontier, so it is deduplicated, checked against that host's robots rules, and fetched by that host's workers under its concurrency limit — a cross-host redirect otherwise borrows the original host's slot to hit a different site. Following redirects automatically is simpler and fine for same-host redirects; a middle ground is to follow automatically but check response.url's host and discard results that left the allowed scope. Record permanent redirects (301, 308) so the old URL is not crawled again.

Verify: a redirect from host A to host B is fetched by B's workers, after B's robots check.

3. Use a total timeout, not just a read timeout

sock_read bounds the gap between received chunks. A server that keeps sending — slowly — never trips it:

CRAWL_TIMEOUT = aiohttp.ClientTimeout(
    total=15,             # the whole request, including redirects and body
    sock_connect=5,       # TCP + TLS handshake
    sock_read=10,         # gap between chunks
)

async with session.get(url, timeout=CRAWL_TIMEOUT) as response:
    body = await read_limited(response, MAX_BODY)

Measured: with only sock_read=2, a server sending one byte every 0.5 s was still being read after 12 s, having delivered 25 bytes; with total=5, the request failed with TimeoutError at 5.3 s. In a crawler, a handful of such servers would hold workers indefinitely, and per-host workers mean the host's whole crawl stalls. The total timeout covers everything inside the request — connection, redirects, headers and body — and is the one that makes each fetch's cost predictable.

Verify: a test endpoint that drips data is abandoned at the total timeout, and the worker moves on.

Time spent on a server sending 1 byte every 0.5 s 2 horizontal bars comparing sock_read=2 only with the others. Time spent on a server sending 1 byte every 0.5 s sock_read=2 only >12 s, still reading total=5 5.3 s, TimeoutError aiohttp 3.14; the sock_read run was abandoned by the test after 12 s. Only a total deadline bounds a slow-but-steady sender.

4. Cap the body size and check the content type

A crawler should not download a 4 GB ISO because a link pointed at it. Check headers first, then read with a limit:

MAX_BODY = 5 * 1024 * 1024
HTML_TYPES = ("text/html", "application/xhtml+xml")


async def read_page(response: aiohttp.ClientResponse) -> str | None:
    ctype = response.headers.get("Content-Type", "").split(";")[0].strip().lower()
    if ctype not in HTML_TYPES:
        return None                                       # not a page: skip without reading
    declared = response.content_length
    if declared is not None and declared > MAX_BODY:
        return None
    data = bytearray()
    async for chunk in response.content.iter_chunked(64 * 1024):
        data += chunk
        if len(data) > MAX_BODY:
            return None                                   # lied about length, or chunked and huge
    return data.decode(response.get_encoding(), errors="replace")

Leaving the async with block without reading the body releases the connection; aiohttp closes it rather than returning a half-read connection to the pool. Content-Length can be absent or wrong, so the streaming check is the real limit. Note that response.content.read(n) returns up to n bytes — often a single chunk — so it is not a way to read a whole body with a cap.

Verify: links to large binaries are skipped without downloading, and an HTML page larger than the cap is abandoned at the cap.

5. Classify failures for retries and reporting

Each failure class needs a different response, and recording them makes a crawl's health visible:

async def fetch_with_policy(session, url: str) -> tuple[str, str | None]:
    try:
        async with session.get(url, timeout=CRAWL_TIMEOUT, max_redirects=5) as response:
            if response.status == 429 or response.status >= 500:
                return "retry_later", None               # back off this host
            if response.status >= 400:
                return "gone", None                       # 404, 410: do not retry
            return "ok", await read_page(response)
    except aiohttp.TooManyRedirects:
        return "redirect_loop", None
    except TimeoutError:
        return "timeout", None
    except aiohttp.ClientConnectorError:
        return "unreachable", None                        # DNS failure, refused
    except aiohttp.ClientError:
        return "protocol_error", None

Retry retry_later and timeout a couple of times with growing delays, and slow the host down; never retry gone or redirect_loop. Count every outcome per host, because a host that returns mostly timeouts or 503s should be paused rather than hammered — the adaptive limit in limiting concurrency per host in an async crawler consumes exactly these signals.

Verify: every fetch ends in exactly one outcome class, and per-host counts are visible during the crawl.

What should happen after this fetch? A decision on How did the fetch end with 4 outcomes. What should happen after this fetch? How did the fetch end? redirect to another host back to frontier that host's rules 429, 5xx, timeout retry later + slow host bounded attempts 404, 410, redirect loop record, no retry gone not HTML or too big skip unread connection released Every fetch is bounded, classified, and fed back into scheduling.

Verification

Fetches are bounded when:

  • Redirects are capped (about 5) and links resolve against response.url.
  • Cross-host redirects re-enter the frontier or are scope-checked.
  • A total timeout bounds every request, not just read gaps.
  • Content type and body size are checked before and while reading.

Diagnostic Hook: histogram fetch duration per host and count outcomes by class. A cluster of fetches at exactly the total timeout identifies slow or dripping servers; many redirect-loop outcomes on one host usually mean a cookie or consent wall the crawler cannot pass.

Pitfalls & edge cases

  • Only sock_read. Measured: a dripping server held a fetch past 12 s.
  • Resolving links against the requested URL. Redirected pages produce wrong links.
  • Automatic cross-host redirects. They bypass the target host's limits and robots rules.
  • content.read(n) as a size limit. It returns at most one chunk, not the body.

Frequently Asked Questions

How many redirects does aiohttp follow?

Ten by default; a redirect loop raised TooManyRedirects after 10 hops in testing. Set max_redirects lower for crawling, around 5.

Why doesn't aiohttp's sock_read timeout stop slow servers?

It bounds the gap between chunks. A server sending a byte every 0.5 s never exceeds a 2 s gap; in testing it was still being read after 12 s. Use ClientTimeout(total=...) to bound the whole request.

How do I limit download size in aiohttp?

Check Content-Type and Content-Length first, then read with iter_chunked and stop once the accumulated size passes your limit.

Should a crawler follow redirects automatically?

Same-host redirects, yes. Cross-host redirects are better sent back through the frontier with allow_redirects=False, so the target host's robots rules and concurrency limits apply.