Handling Redirects and Timeouts When Crawling with aiohttp¶
A crawler fetches from servers it knows nothing about, so every request must be bounded in hops, time and size — and bounded in ways that hostile or broken servers cannot get around. aiohttp provides the controls; their defaults and exact semantics matter. Tested with aiohttp 3.14 against a local test site: a redirect was followed transparently with its 302 recorded in response.history; a URL that redirected to itself raised TooManyRedirects after 10 hops, the default limit; and a server that sent one byte every half-second was still being read after 12 s with a sock_read timeout of 2 s, because each byte arrived well within 2 s. A total timeout of 5 s stopped it at 5.3 s. This guide sets each limit for crawling, decides what to do with redirects across hosts, and caps body size.
Prerequisites¶
- Python 3.11+,
pip install aiohttp; measured with aiohttp 3.14. - aiohttp timeouts, from configuring aiohttp TCPConnector limits.
- Per-host scheduling, from limiting concurrency per host in an async crawler.
1. Bound redirects and record where you landed¶
aiohttp follows redirects by default, up to max_redirects=10, and records each hop:
async with session.get(url, max_redirects=5) as response:
final_url = str(response.url) # where the content actually came from
hops = [(h.status, str(h.url)) for h in response.history]
Tested: /redirect/5 returned the content of /p/5 with history == [302]; a self-redirecting URL raised aiohttp.TooManyRedirects after 10 hops, and after 3 with max_redirects=3. Five is plenty for real sites — chains longer than that are usually loops or tracking hops. Always use response.url, not the requested URL, as the base for resolving relative links in the page and as the URL you record as fetched; otherwise links on redirected pages resolve against the wrong path.
Verify: a redirect loop on a test host fails fast with TooManyRedirects, and links on a redirected page resolve against the final URL.
2. Decide what cross-host redirects mean¶
A redirect can move the request to a different host — http to https, example.com to www.example.com, or a different site entirely. That matters for politeness, robots rules and scope:
async def fetch(session, url: str, scheduler) -> FetchResult | None:
async with session.get(url, allow_redirects=False) as response:
if response.status in (301, 302, 303, 307, 308):
location = response.headers.get("Location")
if location:
target = normalize(urljoin(url, location))
scheduler.offer(target) # re-enters the frontier: dedup, robots, host queue
return None
return await read_page(response)
Handling redirects yourself with allow_redirects=False sends each target back through the frontier, so it is deduplicated, checked against that host's robots rules, and fetched by that host's workers under its concurrency limit — a cross-host redirect otherwise borrows the original host's slot to hit a different site. Following redirects automatically is simpler and fine for same-host redirects; a middle ground is to follow automatically but check response.url's host and discard results that left the allowed scope. Record permanent redirects (301, 308) so the old URL is not crawled again.
Verify: a redirect from host A to host B is fetched by B's workers, after B's robots check.
3. Use a total timeout, not just a read timeout¶
sock_read bounds the gap between received chunks. A server that keeps sending — slowly — never trips it:
CRAWL_TIMEOUT = aiohttp.ClientTimeout(
total=15, # the whole request, including redirects and body
sock_connect=5, # TCP + TLS handshake
sock_read=10, # gap between chunks
)
async with session.get(url, timeout=CRAWL_TIMEOUT) as response:
body = await read_limited(response, MAX_BODY)
Measured: with only sock_read=2, a server sending one byte every 0.5 s was still being read after 12 s, having delivered 25 bytes; with total=5, the request failed with TimeoutError at 5.3 s. In a crawler, a handful of such servers would hold workers indefinitely, and per-host workers mean the host's whole crawl stalls. The total timeout covers everything inside the request — connection, redirects, headers and body — and is the one that makes each fetch's cost predictable.
Verify: a test endpoint that drips data is abandoned at the total timeout, and the worker moves on.
4. Cap the body size and check the content type¶
A crawler should not download a 4 GB ISO because a link pointed at it. Check headers first, then read with a limit:
MAX_BODY = 5 * 1024 * 1024
HTML_TYPES = ("text/html", "application/xhtml+xml")
async def read_page(response: aiohttp.ClientResponse) -> str | None:
ctype = response.headers.get("Content-Type", "").split(";")[0].strip().lower()
if ctype not in HTML_TYPES:
return None # not a page: skip without reading
declared = response.content_length
if declared is not None and declared > MAX_BODY:
return None
data = bytearray()
async for chunk in response.content.iter_chunked(64 * 1024):
data += chunk
if len(data) > MAX_BODY:
return None # lied about length, or chunked and huge
return data.decode(response.get_encoding(), errors="replace")
Leaving the async with block without reading the body releases the connection; aiohttp closes it rather than returning a half-read connection to the pool. Content-Length can be absent or wrong, so the streaming check is the real limit. Note that response.content.read(n) returns up to n bytes — often a single chunk — so it is not a way to read a whole body with a cap.
Verify: links to large binaries are skipped without downloading, and an HTML page larger than the cap is abandoned at the cap.
5. Classify failures for retries and reporting¶
Each failure class needs a different response, and recording them makes a crawl's health visible:
async def fetch_with_policy(session, url: str) -> tuple[str, str | None]:
try:
async with session.get(url, timeout=CRAWL_TIMEOUT, max_redirects=5) as response:
if response.status == 429 or response.status >= 500:
return "retry_later", None # back off this host
if response.status >= 400:
return "gone", None # 404, 410: do not retry
return "ok", await read_page(response)
except aiohttp.TooManyRedirects:
return "redirect_loop", None
except TimeoutError:
return "timeout", None
except aiohttp.ClientConnectorError:
return "unreachable", None # DNS failure, refused
except aiohttp.ClientError:
return "protocol_error", None
Retry retry_later and timeout a couple of times with growing delays, and slow the host down; never retry gone or redirect_loop. Count every outcome per host, because a host that returns mostly timeouts or 503s should be paused rather than hammered — the adaptive limit in limiting concurrency per host in an async crawler consumes exactly these signals.
Verify: every fetch ends in exactly one outcome class, and per-host counts are visible during the crawl.
Verification¶
Fetches are bounded when:
- Redirects are capped (about 5) and links resolve against
response.url. - Cross-host redirects re-enter the frontier or are scope-checked.
- A total timeout bounds every request, not just read gaps.
- Content type and body size are checked before and while reading.
Diagnostic Hook: histogram fetch duration per host and count outcomes by class. A cluster of fetches at exactly the total timeout identifies slow or dripping servers; many redirect-loop outcomes on one host usually mean a cookie or consent wall the crawler cannot pass.
Pitfalls & edge cases¶
- Only
sock_read. Measured: a dripping server held a fetch past 12 s. - Resolving links against the requested URL. Redirected pages produce wrong links.
- Automatic cross-host redirects. They bypass the target host's limits and robots rules.
content.read(n)as a size limit. It returns at most one chunk, not the body.
Frequently Asked Questions¶
How many redirects does aiohttp follow?
Ten by default; a redirect loop raised TooManyRedirects after 10 hops in testing. Set max_redirects lower for crawling, around 5.
Why doesn't aiohttp's sock_read timeout stop slow servers?
It bounds the gap between chunks. A server sending a byte every 0.5 s never exceeds a 2 s gap; in testing it was still being read after 12 s. Use ClientTimeout(total=...) to bound the whole request.
How do I limit download size in aiohttp?
Check Content-Type and Content-Length first, then read with iter_chunked and stop once the accumulated size passes your limit.
Should a crawler follow redirects automatically?
Same-host redirects, yes. Cross-host redirects are better sent back through the frontier with allow_redirects=False, so the target host's robots rules and concurrency limits apply.
Related¶
- Async Web Crawlers — up to the topic overview.
- Respecting robots.txt and crawl-delay asynchronously — the rules a redirect target must also satisfy.
- Network I/O & Protocol Handling — the section overview.