Skip to content

Respecting robots.txt and Crawl-delay Asynchronously

robots.txt is how a site tells crawlers what not to fetch and how fast to go. An async crawler has to fetch it once per host before the first request — even when dozens of requests for that host start at the same moment — cache the parsed rules, check every URL, and pace requests to the host. Measured on a four-host test site: ignoring robots fetched 27 pages under a disallowed path; respecting it fetched 0, with exactly 4 robots requests for 4 hosts. One host declared Crawl-delay: 0.05. Python's urllib.robotparser returned None for it — it parses only whole numbers, tested with 0.05, 1 and 2.5 — so the first version of the crawler hit that host at 359 requests per second. Parsing the value by hand and serving the host with one paced worker brought it to 16 per second. This guide implements the full protocol for asyncio.

Prerequisites

  • Python 3.11+, pip install aiohttp.
  • Per-host scheduling, from limiting concurrency per host in an async crawler.
  • The protocol itself: RFC 9309 (the Robots Exclusion Protocol) defines User-agent, Allow and Disallow; Crawl-delay and Sitemap are widely used extensions.

1. Fetch robots.txt once per host, even under concurrency

The first URLs for a new host often arrive together. Without coordination, each would fetch robots.txt. A per-host lock with a cache inside makes the first caller fetch and the rest wait for its result:

import collections
from urllib.robotparser import RobotFileParser

import aiohttp


class RobotsCache:
    def __init__(self, session: aiohttp.ClientSession, agent: str) -> None:
        self.session, self.agent = session, agent
        self.rules: dict[str, RobotFileParser] = {}
        self.locks: dict[str, asyncio.Lock] = collections.defaultdict(asyncio.Lock)

    async def for_origin(self, origin: str) -> RobotFileParser:
        if origin in self.rules:
            return self.rules[origin]
        async with self.locks[origin]:
            if origin not in self.rules:                     # re-check after waiting
                self.rules[origin] = await self._fetch(origin)
        return self.rules[origin]

Tested: four hosts, four robots requests, no matter how many URLs per host were scheduled at once. The double check inside the lock is what prevents the stampede: callers that waited find the rules already cached. Cache entries should expire — a day is common — so long crawls pick up changes, which the RFC also recommends.

Verify: the access log of a test site shows exactly one robots.txt request per crawl, even with high per-host concurrency.

2. Treat fetch failures the way the protocol says

What robots.txt returns decides what you may crawl. RFC 9309 is specific, and the cautious choices differ for client and server errors:

    async def _fetch(self, origin: str) -> RobotFileParser:
        rp = RobotFileParser()
        try:
            async with self.session.get(f"{origin}/robots.txt", max_redirects=5,
                                        timeout=aiohttp.ClientTimeout(total=10)) as r:
                if r.status >= 500:
                    rp.disallow_all = True           # server error: assume everything disallowed
                elif r.status >= 400:
                    rp.allow_all = True              # 404 etc.: no rules, crawl allowed
                else:
                    raw = bytearray()
                    async for chunk in r.content.iter_chunked(64 * 1024):
                        raw += chunk
                        if len(raw) >= 500 * 1024:           # RFC 9309: parse at least 500 KiB
                            break
                    text = bytes(raw[:500 * 1024]).decode(errors="replace")
                    rp.parse(text.splitlines())
                    rp.fractional_delay = parse_crawl_delay(text, self.agent)
        except (aiohttp.ClientError, TimeoutError):
            rp.disallow_all = True                   # unreachable: back off for now
        return rp

A 404 means the site has no rules: everything is allowed. A 5xx or a network failure means the rules are unknown, so the safe reading is "disallow everything" and retry later. Redirects are followed up to a small limit. Stopping after 500 KiB follows the RFC's minimum parsing limit and protects the crawler from a huge or endless response.

Verify: test hosts that answer robots.txt with 404, 503 and a timeout produce allow-all, disallow-all and disallow-all respectively.

What the robots.txt response means for the crawl A grid of 4 rows by 3 columns. What the robots.txt response means for the crawl robots.txt result rules to apply retry 200 OK parse and apply cache ~24 h 404 / other 4xx allow everything cache ~24 h 5xx disallow everything retry soon timeout / connection error disallow everything retry soon Unknown rules mean back off; missing rules mean go ahead.

3. Check every URL against the rules

Check after normalization and before scheduling, so disallowed URLs never enter a host's queue:

AGENT = "examplebot"


async def admit(url: str, robots: RobotsCache, scheduler) -> bool:
    origin = "{0.scheme}://{0.netloc}".format(urlsplit(url))
    rules = await robots.for_origin(origin)
    if not rules.can_fetch(AGENT, url):
        DISALLOWED.inc()
        return False
    scheduler.submit(url)
    return True

Measured: with the check, none of the 27 URLs under /private/ were fetched; without it, all 27 were. Use the same product token in can_fetch that appears in your User-Agent header, so site rules written for your crawler by name are applied. Checking before scheduling also keeps the per-host queues free of URLs that would only be discarded later.

Verify: a test site with a Disallow rule sees no requests under that path.

4. Parse Crawl-delay yourself and pace the host

urllib.robotparser.RobotFileParser.crawl_delay() returns the value only if it is a whole number:

rp = RobotFileParser()
rp.parse(["User-agent: *", "Crawl-delay: 0.05"])
rp.crawl_delay("examplebot")          # tested: None (also None for 2.5; 1 for "1")


import re

def parse_crawl_delay(text: str, agent: str) -> float | None:
    """Fractional-aware Crawl-delay for the group matching agent, else the * group."""
    groups: dict[str, float] = {}
    current: list[str] = []
    for line in text.splitlines():
        line = line.split("#", 1)[0].strip()
        if not line or ":" not in line:
            continue
        key, value = (part.strip() for part in line.split(":", 1))
        key = key.lower()
        if key == "user-agent":
            current = [value.lower()]
        elif key == "crawl-delay" and current:
            try:
                groups[current[0]] = float(value)
            except ValueError:
                pass
    return groups.get(agent.lower(), groups.get("*"))

Then serve a delayed host with a single worker that sleeps the delay between requests:

async def serve_host(host: str, queue: asyncio.Queue, rules: RobotFileParser) -> None:
    delay = rules.crawl_delay(AGENT) or getattr(rules, "fractional_delay", None) or 0
    workers = 1 if delay else PER_HOST
    ...
    # in the worker loop, after each fetch:
    await asyncio.sleep(delay)

Measured: with the delay ignored, the test host received 359 requests per second; with one worker and a 0.05 s sleep after each 10 ms response, 16 per second — one request per delay plus response time. This simple grouping takes the first agent of each group, which covers common files; a multi-agent group needs a fuller parser. Treat very large delays (minutes) as a sign to crawl that site rarely rather than to hold a worker asleep.

Verify: the request timestamps for a host with Crawl-delay are spaced at least that far apart.

Requests per second to a host declaring Crawl-delay 0.05 2 horizontal bars comparing delay ignored (robotparser returned None) with the others. Requests per second to a host declaring Crawl-delay 0.05 delay ignored (robotparser returned None) 359 req/s delay parsed and applied 16 req/s Local test host answering in 10 ms; per-host queue with 4 workers when no delay, 1 worker with delay. A silently ignored directive is the easiest way to be rude to a site.

5. Use Sitemap entries and re-check periodically

robots.txt often lists sitemaps, which are cheaper than link discovery. And rules change; long crawls must refresh them:

def sitemaps(rules: RobotFileParser) -> list[str]:
    return rules.site_maps() or []                       # Python 3.8+


class ExpiringRobotsCache(RobotsCache):
    TTL = 24 * 3600

    async def for_origin(self, origin: str) -> RobotFileParser:
        rules = self.rules.get(origin)
        if rules is not None and time.monotonic() - rules.fetched_at > self.TTL:
            del self.rules[origin]                       # refetch on next use
        rules = await super().for_origin(origin)
        rules.fetched_at = getattr(rules, "fetched_at", time.monotonic())
        return rules

Seeding a host's queue from its sitemaps reaches its pages with far fewer requests than following links, and sitemap lastmod dates tell a re-crawl which pages changed. A 24-hour cache lifetime is the convention; a host whose robots fetch failed with a 5xx should be retried much sooner, since its rules were never really known.

Verify: a long-running crawl refetches each host's robots.txt about once a day, and picks up a rule change made mid-crawl.

Robots handling for one host A flow of 5 stages. Robots handling for one host first URL for host lock, fetch once map status 200 parse, 4xx allow, 5xx deny parse rules, delay, sitemaps per URL can_fetch before queueing pace + refresh delay, 24 h TTL Fetch once, decide per URL, pace per host.

Verification

The robots protocol is respected when:

  • Each host's robots.txt is fetched once per cache period, even under concurrency.
  • Status codes map to the right rules: 4xx allows, 5xx and errors disallow.
  • Every URL is checked before scheduling with the crawler's own product token.
  • Crawl-delay is honoured, including fractional values, by pacing the host.

Diagnostic Hook: count robots fetches per host per day, disallowed URLs skipped, and the minimum interval between requests to each host with a declared delay. More than one robots fetch per host per period means the lock is missing; a minimum interval below the declared delay means the delay is being ignored — the measured 359 requests per second.

Pitfalls & edge cases

  • Fetching robots.txt per request. One fetch per host per period is enough.
  • Treating a 5xx as "no rules". The RFC says to assume everything is disallowed.
  • Relying on crawl_delay() alone. Fractional values return None.
  • A different name in can_fetch than in User-Agent. Site rules for your crawler are missed.

Frequently Asked Questions

How do I respect robots.txt in an asyncio crawler?

Fetch robots.txt once per host under a per-host lock, parse it with urllib.robotparser, check can_fetch for every URL before queueing it, and pace each host according to its Crawl-delay.

Why does urllib.robotparser return None for Crawl-delay?

It only parses whole-number values. In testing, Crawl-delay: 0.05 and 2.5 both returned None and 1 returned 1. Parse the directive yourself if sites use fractions.

What should a crawler do if robots.txt returns 404 or 500?

Under RFC 9309, a 4xx means there are no rules and crawling is allowed; a 5xx or network failure means the rules are unknown and everything should be treated as disallowed until a later retry succeeds.

How often should a crawler refetch robots.txt?

About once a day, which is the RFC's guidance for caching, and sooner for hosts whose previous fetch failed.