Async Web Crawlers in Python¶
A web crawler is the most demanding I/O workload most Python developers ever write: thousands of concurrent requests to hosts you do not control, a work list that grows as you go, and every failure mode the network offers — slow servers, redirect loops, connections that trickle one byte at a time, sites that ask not to be crawled at all. asyncio is a good fit because the work is almost all waiting. It is also easy to get wrong in ways that are rude to the sites you crawl and wasteful for you. Measured on a local test site of four hosts, one of them slow: a crawler with only a global concurrency limit put all 20 of its requests on the slow host; per-host semaphores in front of one shared queue fixed that and took 16.28 s instead of 3.87 s; per-host queues fetched the same budget in 1.24 s with every host capped at 4 concurrent requests. Without URL normalization, 206 of 1,219 fetches (17%) were duplicates of pages already fetched under a different spelling. Ignoring robots.txt fetched 27 disallowed pages; respecting it fetched none and slowed the crawl-delayed host from 359 to 16 requests per second.
This section covers the parts that make a crawler production-ready: scheduling by host, the robots protocol, the URL frontier and deduplication, HTTP edge cases, and state that survives a crash. The parent section, Network I/O & Protocol Handling, covers the HTTP clients and connection pools a crawler is built on.
Scope of this section:
- Per-host concurrency and queues, so politeness and throughput do not conflict.
robots.txtandCrawl-delay, fetched once per host and honoured per request.- URL normalization and a deduplicating frontier.
- Redirects, timeouts and response limits with aiohttp.
- Persisting crawl state so a crash resumes instead of restarting.
Architectural principles¶
- Politeness is per host. Every limit that protects a site — concurrency, delay, backoff — is scoped to that site, never shared globally.
- Isolation by host. Each host has its own queue and workers, so a slow or failing site only slows itself.
- One canonical form per URL. URLs are normalized before deduplication and scheduling, so the same page is fetched once.
- Every fetch is bounded. Redirect count, total time and body size have limits; a crawler meets every pathological server eventually.
- State outlives the process. The frontier and the set of fetched URLs are durable, so a crash costs the in-flight requests, not the whole crawl.
Execution model: many hosts, few requests each¶
A crawler's concurrency comes from breadth, not depth. One host should see a handful of requests at a time; the crawler as a whole can run hundreds by talking to many hosts at once. That shape decides the data structures. A single shared queue with a global pool of workers lets URLs for one host crowd out all others — and because slow hosts keep requests in flight longest, they end up with most of the workers, which is the opposite of polite. Adding per-host semaphores to that design keeps the per-host counts down but creates head-of-line blocking: workers pick up a URL for a busy host and wait, while URLs for idle hosts sit behind them. The measured 16.28 s against 1.24 s is that blocking.
class Crawler:
def __init__(self, session, per_host: int = 4, max_hosts: int = 100) -> None:
self.session = session
self.per_host = per_host
self.host_slots = asyncio.Semaphore(max_hosts)
self.queues: dict[str, asyncio.Queue[str]] = {}
def schedule(self, url: str) -> None:
host = urlsplit(url).netloc
if host not in self.queues:
self.queues[host] = asyncio.Queue()
asyncio.create_task(self.serve_host(host))
self.queues[host].put_nowait(url)
async def serve_host(self, host: str) -> None:
async with self.host_slots: # bounded number of active hosts
async with asyncio.TaskGroup() as tg:
for _ in range(self.per_host): # bounded requests per host
tg.create_task(self.host_worker(host))
Per-host queues with a few workers each, plus a cap on how many hosts are active, give both properties: each host sees at most per_host requests, and total concurrency is per_host × max_hosts. The details, including adaptive limits that back off on 429s, are in limiting concurrency per host in an async crawler.
Pattern catalogue¶
Fetch robots.txt once per host, and honour Crawl-delay¶
robots.txt is fetched once per host before the first request and cached; concurrent first requests share one fetch through a per-host lock. The standard library's urllib.robotparser parses it, but it ignores fractional Crawl-delay values — tested, 0.05 and 2.5 returned None, only integers parsed — so read that directive yourself. Measured, respecting the rules fetched 4 robots files for 4 hosts, skipped all 27 disallowed URLs, and paced the delayed host at 16 requests per second instead of 359. See respecting robots.txt and crawl-delay asynchronously.
async def robots_for(origin: str) -> RobotFileParser:
async with robots_locks[origin]: # one fetch per host, even under concurrency
if origin not in robots_cache:
robots_cache[origin] = await fetch_and_parse_robots(origin)
return robots_cache[origin]
Normalize before you deduplicate¶
The same page appears under many spellings: with and without a fragment, with tracking parameters, with a trailing slash, with an upper-case scheme or host, with parameters in a different order. Normalizing to one canonical form before checking the seen set removed every duplicate in the test crawl — 206 of 1,219 fetches without it. Normalization must be conservative: lower-case the scheme and host, drop the fragment and default port, remove known tracking parameters, sort the rest; never lower-case paths, which are case-sensitive on most servers. See deduplicating URLs in a concurrent crawl frontier.
def normalize(url: str) -> str:
s = urlsplit(url)
query = urlencode(sorted((k, v) for k, v in parse_qsl(s.query) if k not in TRACKING))
path = s.path.rstrip("/") or "/"
return urlunsplit((s.scheme.lower(), s.netloc.lower(), path, query, ""))
Bound redirects, time and size on every fetch¶
aiohttp follows up to 10 redirects by default and raised TooManyRedirects on a redirect loop after 10 hops; lower it for crawling. Its sock_read timeout bounds the gap between chunks, so a server that sent one byte every half-second was still being read after 12 s with sock_read=2; a total timeout of 5 s stopped it at 5.3 s. Read bodies in chunks and stop at a size limit. See handling redirects and timeouts when crawling with aiohttp.
timeout = aiohttp.ClientTimeout(total=15, sock_connect=5, sock_read=10)
async with session.get(url, timeout=timeout, max_redirects=5) as response:
body = await read_limited(response, max_bytes=5 * 1024 * 1024)
Keep the frontier durable¶
Store every discovered URL with a state — queued, in flight, done — in a local database, and mark a URL done in the same step that records its outgoing links. On restart, requeue anything left in flight. Measured with SQLite: a crawl killed with SIGKILL after 519 fetches resumed and finished at 607 total fetches with 0 duplicates; a crawler without state would have repeated all 519. See persisting crawl state to resume after a crash.
await db.execute("update urls set state = 0 where state = 1") # in flight at crash -> requeue
Choosing the crawler's shape¶
The right shape depends on how many hosts and how much data:
For a single site, a small fixed concurrency (2–4), robots checks and an in-memory seen set are enough. A crawl across a known list of sites needs the per-host structure. An open-ended crawl that discovers hosts as it goes needs a durable frontier, a cap on active hosts, adaptive per-host limits and good logging of every failure class. Very large crawls outgrow one process: the frontier moves to a shared store and hosts are partitioned across workers so each host is still served by exactly one scheduler.
Resource boundaries¶
A crawler meets every resource limit at once, so set each explicitly:
- Sockets and file descriptors:
per_host × max_active_hosts, plus DNS and robots fetches; raiseulimit -naccordingly and set the connector'slimitto match. - Memory: the seen set grows with every discovered URL — roughly 100–200 bytes per URL in a Python set of strings — so tens of millions of URLs belong in a database or a probabilistic filter, not a set.
- Body size: cap every response; a single multi-gigabyte download should not take the crawler down.
- DNS: cache answers (
ttl_dns_cache), since a broad crawl resolves thousands of names; see caching DNS lookups in async HTTP clients. - CPU: HTML parsing is CPU-bound; at high page rates, parse in a process pool so the event loop keeps fetching.
Integrated production example¶
A compact crawler that combines the patterns — per-host scheduling, robots, normalization and bounded fetches — with an in-memory frontier for clarity:
import asyncio
import collections
import re
from urllib.parse import parse_qsl, urlencode, urljoin, urlsplit, urlunsplit
from urllib.robotparser import RobotFileParser
import aiohttp
HREF = re.compile(r'href="([^"#]+)')
TRACKING = {"utm_source", "utm_medium", "utm_campaign", "gclid", "fbclid"}
AGENT = "examplebot/1.0 (+https://example.com/bot)"
def normalize(url: str) -> str:
s = urlsplit(url)
query = urlencode(sorted((k, v) for k, v in parse_qsl(s.query) if k not in TRACKING))
path = s.path if s.path in ("", "/") else s.path.rstrip("/")
return urlunsplit((s.scheme.lower(), s.netloc.lower(), path or "/", query, ""))
async def read_limited(response: aiohttp.ClientResponse, max_bytes: int) -> bytes:
chunks, size = [], 0
async for chunk in response.content.iter_chunked(64 * 1024):
size += len(chunk)
if size > max_bytes:
raise aiohttp.ClientPayloadError(f"body larger than {max_bytes} bytes")
chunks.append(chunk)
return b"".join(chunks)
class Crawler:
def __init__(self, session: aiohttp.ClientSession, budget: int, per_host: int = 4) -> None:
self.session, self.budget, self.per_host = session, budget, per_host
self.seen: set[str] = set()
self.queues: dict[str, asyncio.Queue[str]] = {}
self.robots: dict[str, RobotFileParser] = {}
self.robots_locks = collections.defaultdict(asyncio.Lock)
self.fetched = 0
self.pending = 0
self.idle = asyncio.Event()
self.tasks: set[asyncio.Task] = set()
def add(self, url: str) -> None:
url = normalize(url)
if url in self.seen or not url.startswith(("http://", "https://")):
return
self.seen.add(url)
host = urlsplit(url).netloc
if host not in self.queues:
self.queues[host] = asyncio.Queue()
for _ in range(self.per_host):
task = asyncio.create_task(self.worker(host))
self.tasks.add(task)
self.pending += 1
self.idle.clear()
self.queues[host].put_nowait(url)
async def rules(self, host: str, scheme: str) -> RobotFileParser:
async with self.robots_locks[host]:
if host not in self.robots:
rp = RobotFileParser()
try:
async with self.session.get(f"{scheme}://{host}/robots.txt",
timeout=aiohttp.ClientTimeout(total=10)) as r:
if r.status >= 500:
rp.disallow_all = True
elif r.status >= 400:
rp.allow_all = True
else:
rp.parse((await r.text()).splitlines())
except aiohttp.ClientError:
rp.allow_all = True
self.robots[host] = rp
return self.robots[host]
async def worker(self, host: str) -> None:
queue = self.queues[host]
while True:
url = await queue.get()
try:
if self.fetched >= self.budget:
continue
rp = await self.rules(host, urlsplit(url).scheme)
if not rp.can_fetch(AGENT, url):
continue
async with self.session.get(url, max_redirects=5,
timeout=aiohttp.ClientTimeout(total=15)) as r:
if "html" not in r.headers.get("Content-Type", ""):
continue
body = (await read_limited(r, 5 * 1024 * 1024)).decode(errors="replace")
self.fetched += 1
for href in HREF.findall(body):
self.add(urljoin(str(r.url), href))
await asyncio.sleep(rp.crawl_delay(AGENT) or 0)
except (aiohttp.ClientError, TimeoutError) as exc:
log.info("failed %s: %r", url, exc)
finally:
self.pending -= 1
if self.pending == 0:
self.idle.set()
async def run(self, seeds: list[str]) -> int:
for seed in seeds:
self.add(seed)
await self.idle.wait()
for task in self.tasks:
task.cancel()
return self.fetched
async def main() -> None:
connector = aiohttp.TCPConnector(limit=400, limit_per_host=4, ttl_dns_cache=600)
async with aiohttp.ClientSession(connector=connector, headers={"User-Agent": AGENT}) as session:
print(await Crawler(session, budget=10_000).run(["https://example.com/"]))
Every request carries an identifying User-Agent with a contact URL, which is basic courtesy and often required by site operators. The production version replaces seen and the queues with the durable frontier and adds the fractional crawl-delay parsing and adaptive limits described in the guides.
Operating a crawler responsibly¶
The technical patterns above are what keep a crawler from harming the sites it visits; a few operational habits complete the picture.
- Identify yourself. Send a
User-Agentthat names the crawler and links to a page explaining what it does and how to contact you or opt out. Operators who can identify a crawler can ask it to slow down; anonymous traffic simply gets blocked. - Scope the crawl. Decide up front which hosts, paths and depths are in scope, and enforce it when URLs are added to the frontier rather than when they are fetched, so out-of-scope URLs never consume memory. An open-ended crawl without scope rules grows until it runs out of disk.
- Prefer sitemaps and feeds. Many sites publish
sitemap.xml(often listed inrobots.txt) and RSS or Atom feeds; reading them finds pages with fewer requests than following every link, and tells you which pages changed. - Re-crawl conditionally. Store
ETagandLast-Modifiedvalues and sendIf-None-MatchandIf-Modified-Sinceon re-crawls; a304 Not Modifiedresponse costs the site almost nothing. - Honour opt-outs quickly. Treat a
403, a429withRetry-After, or a contact request as signals to back off for that host, and keep a deny list that the scheduler checks before scheduling anything. - Respect terms and law. Some sites forbid automated access in their terms, and data protection rules apply to personal data you collect. That is a policy decision outside the code, but the code should make it enforceable through the scope and deny lists.
None of this slows a well-built crawler noticeably: the time goes to waiting on many hosts in parallel, and courtesy towards each one costs only a little latency on that host.
Diagnostic hook callout¶
Diagnostic Hook: export, per host, in-flight requests, fetch rate, error rate by class (timeout, 4xx, 5xx, redirect limit, robots-disallowed), and queue depth; globally, frontier size, seen-set size and duplicate rate. In-flight piling up on one host means scheduling is not per host; workers idle while queues are full means head-of-line blocking; a rising duplicate rate means a new URL variant is escaping normalization; and a frontier that grows faster than the fetch rate forever means the crawl needs a scope rule.
Failure modes¶
- Global limits only. A slow host absorbed all 20 requests in testing.
- Shared queue with per-host semaphores. Head-of-line blocking: 16.28 s against 1.24 s.
- No normalization. 17% of fetches were duplicates in the test crawl.
- Trusting urllib.robotparser for Crawl-delay. Fractional values are silently ignored.
- Only a read timeout. A dripping server held a fetch past 12 s; use a total timeout.
- In-memory state for long crawls. A crash repeats everything.
Frequently Asked Questions¶
Is asyncio good for web crawling in Python?
Yes. Crawling is mostly waiting on many independent hosts, which asyncio handles with one thread. Structure it per host — a queue and a few workers per site — so it is polite and a slow site does not slow the rest.
How many concurrent requests should a crawler make?
A few per host, such as 2 to 4 and fewer if the site asks for a crawl delay, and as many hosts in parallel as your sockets and bandwidth allow. Total concurrency comes from breadth.
How do I avoid crawling the same page twice?
Normalize every URL to a canonical form — lower-case scheme and host, no fragment, no tracking parameters, sorted query — before checking a seen set. In testing, normalization removed 206 duplicate fetches out of 1,219.
Does Python's robotparser support Crawl-delay?
It parses only integer values; fractional delays such as 0.5 return None. Parse the directive yourself if sites use fractional values.
How do I make a crawler resumable?
Keep the frontier in a database with a state per URL, mark URLs done together with their discovered links, and requeue in-flight URLs on start. A killed test crawl resumed with no duplicate fetches.
Related¶
- Limiting concurrency per host in an async crawler — per-host queues, host caps and adaptive limits.
- Respecting robots.txt and crawl-delay asynchronously — fetching, caching and honouring the robots protocol.
- Deduplicating URLs in a concurrent crawl frontier — normalization and seen sets at scale.
- Handling redirects and timeouts when crawling with aiohttp — bounding every fetch.
- Persisting crawl state to resume after a crash — a durable frontier.
- Downloading many URLs concurrently with progress — the simpler case of a fixed URL list.
- Network I/O & Protocol Handling — up to the section overview.