Respecting robots.txt and Crawl-delay Asynchronously¶
robots.txt is how a site tells crawlers what not to fetch and how fast to go. An async crawler has to fetch it once per host before the first request — even when dozens of requests for that host start at the same moment — cache the parsed rules, check every URL, and pace requests to the host. Measured on a four-host test site: ignoring robots fetched 27 pages under a disallowed path; respecting it fetched 0, with exactly 4 robots requests for 4 hosts. One host declared Crawl-delay: 0.05. Python's urllib.robotparser returned None for it — it parses only whole numbers, tested with 0.05, 1 and 2.5 — so the first version of the crawler hit that host at 359 requests per second. Parsing the value by hand and serving the host with one paced worker brought it to 16 per second. This guide implements the full protocol for asyncio.
Prerequisites¶
- Python 3.11+,
pip install aiohttp. - Per-host scheduling, from limiting concurrency per host in an async crawler.
- The protocol itself: RFC 9309 (the Robots Exclusion Protocol) defines
User-agent,AllowandDisallow;Crawl-delayandSitemapare widely used extensions.
1. Fetch robots.txt once per host, even under concurrency¶
The first URLs for a new host often arrive together. Without coordination, each would fetch robots.txt. A per-host lock with a cache inside makes the first caller fetch and the rest wait for its result:
import collections
from urllib.robotparser import RobotFileParser
import aiohttp
class RobotsCache:
def __init__(self, session: aiohttp.ClientSession, agent: str) -> None:
self.session, self.agent = session, agent
self.rules: dict[str, RobotFileParser] = {}
self.locks: dict[str, asyncio.Lock] = collections.defaultdict(asyncio.Lock)
async def for_origin(self, origin: str) -> RobotFileParser:
if origin in self.rules:
return self.rules[origin]
async with self.locks[origin]:
if origin not in self.rules: # re-check after waiting
self.rules[origin] = await self._fetch(origin)
return self.rules[origin]
Tested: four hosts, four robots requests, no matter how many URLs per host were scheduled at once. The double check inside the lock is what prevents the stampede: callers that waited find the rules already cached. Cache entries should expire — a day is common — so long crawls pick up changes, which the RFC also recommends.
Verify: the access log of a test site shows exactly one robots.txt request per crawl, even with high per-host concurrency.
2. Treat fetch failures the way the protocol says¶
What robots.txt returns decides what you may crawl. RFC 9309 is specific, and the cautious choices differ for client and server errors:
async def _fetch(self, origin: str) -> RobotFileParser:
rp = RobotFileParser()
try:
async with self.session.get(f"{origin}/robots.txt", max_redirects=5,
timeout=aiohttp.ClientTimeout(total=10)) as r:
if r.status >= 500:
rp.disallow_all = True # server error: assume everything disallowed
elif r.status >= 400:
rp.allow_all = True # 404 etc.: no rules, crawl allowed
else:
raw = bytearray()
async for chunk in r.content.iter_chunked(64 * 1024):
raw += chunk
if len(raw) >= 500 * 1024: # RFC 9309: parse at least 500 KiB
break
text = bytes(raw[:500 * 1024]).decode(errors="replace")
rp.parse(text.splitlines())
rp.fractional_delay = parse_crawl_delay(text, self.agent)
except (aiohttp.ClientError, TimeoutError):
rp.disallow_all = True # unreachable: back off for now
return rp
A 404 means the site has no rules: everything is allowed. A 5xx or a network failure means the rules are unknown, so the safe reading is "disallow everything" and retry later. Redirects are followed up to a small limit. Stopping after 500 KiB follows the RFC's minimum parsing limit and protects the crawler from a huge or endless response.
Verify: test hosts that answer robots.txt with 404, 503 and a timeout produce allow-all, disallow-all and disallow-all respectively.
3. Check every URL against the rules¶
Check after normalization and before scheduling, so disallowed URLs never enter a host's queue:
AGENT = "examplebot"
async def admit(url: str, robots: RobotsCache, scheduler) -> bool:
origin = "{0.scheme}://{0.netloc}".format(urlsplit(url))
rules = await robots.for_origin(origin)
if not rules.can_fetch(AGENT, url):
DISALLOWED.inc()
return False
scheduler.submit(url)
return True
Measured: with the check, none of the 27 URLs under /private/ were fetched; without it, all 27 were. Use the same product token in can_fetch that appears in your User-Agent header, so site rules written for your crawler by name are applied. Checking before scheduling also keeps the per-host queues free of URLs that would only be discarded later.
Verify: a test site with a Disallow rule sees no requests under that path.
4. Parse Crawl-delay yourself and pace the host¶
urllib.robotparser.RobotFileParser.crawl_delay() returns the value only if it is a whole number:
rp = RobotFileParser()
rp.parse(["User-agent: *", "Crawl-delay: 0.05"])
rp.crawl_delay("examplebot") # tested: None (also None for 2.5; 1 for "1")
import re
def parse_crawl_delay(text: str, agent: str) -> float | None:
"""Fractional-aware Crawl-delay for the group matching agent, else the * group."""
groups: dict[str, float] = {}
current: list[str] = []
for line in text.splitlines():
line = line.split("#", 1)[0].strip()
if not line or ":" not in line:
continue
key, value = (part.strip() for part in line.split(":", 1))
key = key.lower()
if key == "user-agent":
current = [value.lower()]
elif key == "crawl-delay" and current:
try:
groups[current[0]] = float(value)
except ValueError:
pass
return groups.get(agent.lower(), groups.get("*"))
Then serve a delayed host with a single worker that sleeps the delay between requests:
async def serve_host(host: str, queue: asyncio.Queue, rules: RobotFileParser) -> None:
delay = rules.crawl_delay(AGENT) or getattr(rules, "fractional_delay", None) or 0
workers = 1 if delay else PER_HOST
...
# in the worker loop, after each fetch:
await asyncio.sleep(delay)
Measured: with the delay ignored, the test host received 359 requests per second; with one worker and a 0.05 s sleep after each 10 ms response, 16 per second — one request per delay plus response time. This simple grouping takes the first agent of each group, which covers common files; a multi-agent group needs a fuller parser. Treat very large delays (minutes) as a sign to crawl that site rarely rather than to hold a worker asleep.
Verify: the request timestamps for a host with Crawl-delay are spaced at least that far apart.
5. Use Sitemap entries and re-check periodically¶
robots.txt often lists sitemaps, which are cheaper than link discovery. And rules change; long crawls must refresh them:
def sitemaps(rules: RobotFileParser) -> list[str]:
return rules.site_maps() or [] # Python 3.8+
class ExpiringRobotsCache(RobotsCache):
TTL = 24 * 3600
async def for_origin(self, origin: str) -> RobotFileParser:
rules = self.rules.get(origin)
if rules is not None and time.monotonic() - rules.fetched_at > self.TTL:
del self.rules[origin] # refetch on next use
rules = await super().for_origin(origin)
rules.fetched_at = getattr(rules, "fetched_at", time.monotonic())
return rules
Seeding a host's queue from its sitemaps reaches its pages with far fewer requests than following links, and sitemap lastmod dates tell a re-crawl which pages changed. A 24-hour cache lifetime is the convention; a host whose robots fetch failed with a 5xx should be retried much sooner, since its rules were never really known.
Verify: a long-running crawl refetches each host's robots.txt about once a day, and picks up a rule change made mid-crawl.
Verification¶
The robots protocol is respected when:
- Each host's
robots.txtis fetched once per cache period, even under concurrency. - Status codes map to the right rules: 4xx allows, 5xx and errors disallow.
- Every URL is checked before scheduling with the crawler's own product token.
- Crawl-delay is honoured, including fractional values, by pacing the host.
Diagnostic Hook: count robots fetches per host per day, disallowed URLs skipped, and the minimum interval between requests to each host with a declared delay. More than one robots fetch per host per period means the lock is missing; a minimum interval below the declared delay means the delay is being ignored — the measured 359 requests per second.
Pitfalls & edge cases¶
- Fetching robots.txt per request. One fetch per host per period is enough.
- Treating a 5xx as "no rules". The RFC says to assume everything is disallowed.
- Relying on
crawl_delay()alone. Fractional values returnNone. - A different name in
can_fetchthan inUser-Agent. Site rules for your crawler are missed.
Frequently Asked Questions¶
How do I respect robots.txt in an asyncio crawler?
Fetch robots.txt once per host under a per-host lock, parse it with urllib.robotparser, check can_fetch for every URL before queueing it, and pace each host according to its Crawl-delay.
Why does urllib.robotparser return None for Crawl-delay?
It only parses whole-number values. In testing, Crawl-delay: 0.05 and 2.5 both returned None and 1 returned 1. Parse the directive yourself if sites use fractions.
What should a crawler do if robots.txt returns 404 or 500?
Under RFC 9309, a 4xx means there are no rules and crawling is allowed; a 5xx or network failure means the rules are unknown and everything should be treated as disallowed until a later retry succeeds.
How often should a crawler refetch robots.txt?
About once a day, which is the RFC's guidance for caching, and sooner for hosts whose previous fetch failed.
Related¶
- Async Web Crawlers — up to the topic overview.
- Handling redirects and timeouts when crawling with aiohttp — the same bounds apply to robots fetches.
- Network I/O & Protocol Handling — the section overview.