Skip to content

Limiting Concurrency per Host in an Async Crawler

A crawler's global concurrency limit protects your machine; it does nothing for the sites you crawl. With a single limit of 20, nothing stops all 20 requests from landing on one host — and they will, on whichever host is slowest, because its requests stay in flight longest. Measured with an asyncio crawler against four local test hosts, one of them answering in 200 ms instead of 10 ms: a global limit of 20 put 20 concurrent requests on the slow host and fetched 1,219 pages in 3.87 s. Adding a per-host semaphore of 4 in front of a single shared queue capped every host at 4 — and took 16.28 s, because workers blocked on the slow host's semaphore while URLs for fast hosts waited behind them. Giving each host its own queue and four workers fetched the same budget in 1.24 s with every host capped at 4. This guide builds the per-host structure that is both polite and fast.

Prerequisites

1. See where a global limit sends the load

A single queue and a global pool of workers send requests wherever the URLs point. Slow hosts accumulate the in-flight requests:

queue: asyncio.Queue[str] = asyncio.Queue()


async def worker(session: aiohttp.ClientSession) -> None:
    while True:
        url = await queue.get()
        try:
            async with session.get(url) as response:
                await handle(url, await response.text())
        finally:
            queue.task_done()

# 20 workers: measured peak of 20 concurrent requests on the 200 ms host

Measured: the three fast hosts saw at most 4–6 concurrent requests and the slow host saw all 20. That is precisely the wrong way round for politeness — the struggling host gets the most pressure — and it is how crawlers get rate-limited, blocked, or reported. It also makes the crawl's speed depend on the slowest site.

Verify: log concurrent in-flight requests per host; under a global limit, the slowest host's count approaches the whole limit.

Time to fetch ~1,200 pages from four hosts, one slow 3 horizontal bars comparing global limit 20 with the others. Time to fetch ~1,200 pages from four hosts, one slow global limit 20 3.87 s, 20 on slow host shared queue + per-host semaphore 4 16.28 s, head-of-line blocking queue + 4 workers per host 1.24 s, max 4 per host aiohttp 3.14 against four local hosts (10 ms, 10 ms, 10 ms, 200 ms per page); crawl budget 1,200 pages. Per-host limits are only fast when each host also has its own queue.

2. Avoid head-of-line blocking from per-host semaphores

The obvious fix — a semaphore per host inside the same workers — makes the crawl polite and slow:

host_slots: dict[str, asyncio.Semaphore] = collections.defaultdict(lambda: asyncio.Semaphore(4))


async def worker(session) -> None:
    while True:
        url = await queue.get()
        try:
            async with host_slots[urlsplit(url).netloc]:     # a worker can wait here for a long time
                async with session.get(url) as response:
                    await handle(url, await response.text())
        finally:
            queue.task_done()

# measured: every host capped at 4, but the crawl took 16.28 s instead of 3.87 s

When the next URL in the shared queue belongs to the slow host, the worker that took it waits for that host's semaphore. Soon most workers are waiting on the slow host while fast-host URLs sit in the queue behind them — head-of-line blocking. One subtle bug to avoid in this kind of code: an empty defaultdict is falsy, so a check like if host_slots: silently skips the limit until the first host has been seen.

Verify: with per-host semaphores and one queue, most workers are blocked on one host's semaphore whenever that host is slow.

3. Give each host its own queue and workers

Route URLs to a queue per host, and run a small, fixed set of workers per host. A slow host then only slows its own workers:

class HostScheduler:
    def __init__(self, session: aiohttp.ClientSession, per_host: int = 4) -> None:
        self.session = session
        self.per_host = per_host
        self.queues: dict[str, asyncio.Queue[str]] = {}
        self.workers: dict[str, list[asyncio.Task]] = {}
        self.pending = 0
        self.idle = asyncio.Event()

    def submit(self, url: str) -> None:
        host = urlsplit(url).netloc
        if host not in self.queues:
            self.queues[host] = asyncio.Queue()
            self.workers[host] = [asyncio.create_task(self._work(host)) for _ in range(self.per_host)]
        self.pending += 1
        self.idle.clear()
        self.queues[host].put_nowait(url)

    async def _work(self, host: str) -> None:
        queue = self.queues[host]
        while True:
            url = await queue.get()
            try:
                async with self.session.get(url) as response:
                    await handle(url, await response.text())
            except aiohttp.ClientError as exc:
                log.info("fetch failed %s: %s", url, exc)
            finally:
                self.pending -= 1
                if self.pending == 0:
                    self.idle.set()

Measured: 1.24 s for the same crawl budget, with every host capped at 4 concurrent requests. Workers per host bound politeness; the number of hosts with active workers bounds total concurrency — cap that too (step 4). Hosts whose queues stay empty for a while should have their workers retired so idle tasks do not accumulate across a long crawl.

Verify: per-host in-flight never exceeds per_host, and a slow host does not change the fetch rate of the others.

Routing URLs to per-host queues A flow of 5 stages. Routing URLs to per-host queues discovered URL normalize route by host host's own queue N workers per host polite per site global host cap bounded total slow host only slows itself Isolation by host turns a slow site into a local problem.

4. Bound the total and the connections

Per-host workers multiply with the number of hosts. A broad crawl discovers thousands of hosts, so cap how many are active and keep the connection pool consistent with the scheduler:

MAX_ACTIVE_HOSTS = 100
PER_HOST = 4

connector = aiohttp.TCPConnector(
    limit=MAX_ACTIVE_HOSTS * PER_HOST,      # total sockets
    limit_per_host=PER_HOST,                # enforced by the pool too, as a backstop
    ttl_dns_cache=600,
)
active_hosts = asyncio.Semaphore(MAX_ACTIVE_HOSTS)


async def host_session(host: str, queue: asyncio.Queue) -> None:
    async with active_hosts:                # a host's workers run only while it holds a slot
        await drain_host_queue(host, queue)

Setting limit_per_host on the connector equal to the scheduler's per-host worker count means even a bug in the scheduler cannot exceed the politeness limit. A slot per active host also lets the crawler work through hosts in batches when it discovers more than it can serve at once. Size MAX_ACTIVE_HOSTS × PER_HOST from your bandwidth and file-descriptor budget.

Verify: total open sockets stay below the connector limit during a broad crawl.

5. Adapt the per-host limit to the host

A fixed limit of 4 is too much for a small blog and too little for a large CDN-backed site. Adjust per host from what the host tells you:

class AdaptiveHost:
    def __init__(self, start: int = 2, floor: int = 1, ceiling: int = 8) -> None:
        self.limit, self.floor, self.ceiling = start, floor, ceiling

    def on_response(self, status: int, latency: float) -> None:
        if status in (429, 503) or latency > 5.0:
            self.limit = max(self.floor, self.limit // 2)         # back off fast
        elif status < 400 and latency < 0.5:
            self.limit = min(self.ceiling, self.limit + 1)        # grow slowly

Halving on 429, 503 or slow responses and growing by one on fast successes is the additive-increase, multiplicative-decrease rule, and it converges on what each host tolerates. Honour Retry-After when present, and combine the limit with the host's Crawl-delay from respecting robots.txt and crawl-delay asynchronously. The worker count for the host then follows limit — start or retire workers as it changes.

Verify: against a test host that returns 429 above three concurrent requests, the crawler settles at three without sustained errors.

How should this crawler limit concurrency? A decision on How many hosts, and how long with 4 outcomes. How should this crawler limit concurrency? How many hosts, and how long? one host single small limit e.g. 2-4 a few hosts queue + workers per host 1.24 s vs 16.28 s thousands of hosts cap active hosts + connector bounded sockets long-running crawl adaptive per-host limit AIMD on 429/latency Politeness is per host; throughput comes from many hosts in parallel.

Verification

Per-host limiting is correct when:

  • No host ever sees more than its limit of concurrent requests.
  • A slow host does not slow the others, because each host has its own queue.
  • Active hosts and total sockets are capped, consistently with the connector.
  • Limits adapt to 429s, 503s and latency for long crawls.

Diagnostic Hook: chart in-flight requests and fetch rate per host. In-flight piling up on one host under a global limit means the per-host structure is missing; workers blocked on semaphores while other queues are non-empty means a shared queue is causing head-of-line blocking.

Pitfalls & edge cases

  • Only a global limit. Measured: the slow host received all 20 requests.
  • Per-host semaphores with one shared queue. Measured: 16.28 s against 1.24 s.
  • Truthiness checks on defaultdicts. An empty one is falsy and skips the limit.
  • Unbounded hosts. Workers per host multiply with discovered hosts.

Frequently Asked Questions

How do I limit requests per host in an asyncio crawler?

Give each host its own queue and a fixed number of workers that only fetch from it, and set aiohttp's limit_per_host to the same number as a backstop. In testing that crawled in 1.24 s with every host capped at 4.

Why is my crawler slow with per-host semaphores?

With one shared queue, workers take a URL for a slow host and wait on its semaphore while URLs for other hosts wait behind them. In testing that took 16.28 s instead of 1.24 s with per-host queues.

What concurrency should a crawler use per site?

Start low, around 2 to 4, respect Crawl-delay, and adapt: halve on 429, 503 or slow responses and increase slowly on fast successes.

Is aiohttp's limit_per_host enough for a polite crawler?

It caps connections per host, but with a shared work queue it causes the same head-of-line blocking. Use it as a backstop behind per-host queues.