Limiting Concurrency per Host in an Async Crawler¶
A crawler's global concurrency limit protects your machine; it does nothing for the sites you crawl. With a single limit of 20, nothing stops all 20 requests from landing on one host — and they will, on whichever host is slowest, because its requests stay in flight longest. Measured with an asyncio crawler against four local test hosts, one of them answering in 200 ms instead of 10 ms: a global limit of 20 put 20 concurrent requests on the slow host and fetched 1,219 pages in 3.87 s. Adding a per-host semaphore of 4 in front of a single shared queue capped every host at 4 — and took 16.28 s, because workers blocked on the slow host's semaphore while URLs for fast hosts waited behind them. Giving each host its own queue and four workers fetched the same budget in 1.24 s with every host capped at 4. This guide builds the per-host structure that is both polite and fast.
Prerequisites¶
- Python 3.11+,
pip install aiohttp; the measurements use a local four-host test site. - Worker pools, from building an async worker pool with TaskGroup.
- Connector limits, from configuring aiohttp TCPConnector limits.
1. See where a global limit sends the load¶
A single queue and a global pool of workers send requests wherever the URLs point. Slow hosts accumulate the in-flight requests:
queue: asyncio.Queue[str] = asyncio.Queue()
async def worker(session: aiohttp.ClientSession) -> None:
while True:
url = await queue.get()
try:
async with session.get(url) as response:
await handle(url, await response.text())
finally:
queue.task_done()
# 20 workers: measured peak of 20 concurrent requests on the 200 ms host
Measured: the three fast hosts saw at most 4–6 concurrent requests and the slow host saw all 20. That is precisely the wrong way round for politeness — the struggling host gets the most pressure — and it is how crawlers get rate-limited, blocked, or reported. It also makes the crawl's speed depend on the slowest site.
Verify: log concurrent in-flight requests per host; under a global limit, the slowest host's count approaches the whole limit.
2. Avoid head-of-line blocking from per-host semaphores¶
The obvious fix — a semaphore per host inside the same workers — makes the crawl polite and slow:
host_slots: dict[str, asyncio.Semaphore] = collections.defaultdict(lambda: asyncio.Semaphore(4))
async def worker(session) -> None:
while True:
url = await queue.get()
try:
async with host_slots[urlsplit(url).netloc]: # a worker can wait here for a long time
async with session.get(url) as response:
await handle(url, await response.text())
finally:
queue.task_done()
# measured: every host capped at 4, but the crawl took 16.28 s instead of 3.87 s
When the next URL in the shared queue belongs to the slow host, the worker that took it waits for that host's semaphore. Soon most workers are waiting on the slow host while fast-host URLs sit in the queue behind them — head-of-line blocking. One subtle bug to avoid in this kind of code: an empty defaultdict is falsy, so a check like if host_slots: silently skips the limit until the first host has been seen.
Verify: with per-host semaphores and one queue, most workers are blocked on one host's semaphore whenever that host is slow.
3. Give each host its own queue and workers¶
Route URLs to a queue per host, and run a small, fixed set of workers per host. A slow host then only slows its own workers:
class HostScheduler:
def __init__(self, session: aiohttp.ClientSession, per_host: int = 4) -> None:
self.session = session
self.per_host = per_host
self.queues: dict[str, asyncio.Queue[str]] = {}
self.workers: dict[str, list[asyncio.Task]] = {}
self.pending = 0
self.idle = asyncio.Event()
def submit(self, url: str) -> None:
host = urlsplit(url).netloc
if host not in self.queues:
self.queues[host] = asyncio.Queue()
self.workers[host] = [asyncio.create_task(self._work(host)) for _ in range(self.per_host)]
self.pending += 1
self.idle.clear()
self.queues[host].put_nowait(url)
async def _work(self, host: str) -> None:
queue = self.queues[host]
while True:
url = await queue.get()
try:
async with self.session.get(url) as response:
await handle(url, await response.text())
except aiohttp.ClientError as exc:
log.info("fetch failed %s: %s", url, exc)
finally:
self.pending -= 1
if self.pending == 0:
self.idle.set()
Measured: 1.24 s for the same crawl budget, with every host capped at 4 concurrent requests. Workers per host bound politeness; the number of hosts with active workers bounds total concurrency — cap that too (step 4). Hosts whose queues stay empty for a while should have their workers retired so idle tasks do not accumulate across a long crawl.
Verify: per-host in-flight never exceeds per_host, and a slow host does not change the fetch rate of the others.
4. Bound the total and the connections¶
Per-host workers multiply with the number of hosts. A broad crawl discovers thousands of hosts, so cap how many are active and keep the connection pool consistent with the scheduler:
MAX_ACTIVE_HOSTS = 100
PER_HOST = 4
connector = aiohttp.TCPConnector(
limit=MAX_ACTIVE_HOSTS * PER_HOST, # total sockets
limit_per_host=PER_HOST, # enforced by the pool too, as a backstop
ttl_dns_cache=600,
)
active_hosts = asyncio.Semaphore(MAX_ACTIVE_HOSTS)
async def host_session(host: str, queue: asyncio.Queue) -> None:
async with active_hosts: # a host's workers run only while it holds a slot
await drain_host_queue(host, queue)
Setting limit_per_host on the connector equal to the scheduler's per-host worker count means even a bug in the scheduler cannot exceed the politeness limit. A slot per active host also lets the crawler work through hosts in batches when it discovers more than it can serve at once. Size MAX_ACTIVE_HOSTS × PER_HOST from your bandwidth and file-descriptor budget.
Verify: total open sockets stay below the connector limit during a broad crawl.
5. Adapt the per-host limit to the host¶
A fixed limit of 4 is too much for a small blog and too little for a large CDN-backed site. Adjust per host from what the host tells you:
class AdaptiveHost:
def __init__(self, start: int = 2, floor: int = 1, ceiling: int = 8) -> None:
self.limit, self.floor, self.ceiling = start, floor, ceiling
def on_response(self, status: int, latency: float) -> None:
if status in (429, 503) or latency > 5.0:
self.limit = max(self.floor, self.limit // 2) # back off fast
elif status < 400 and latency < 0.5:
self.limit = min(self.ceiling, self.limit + 1) # grow slowly
Halving on 429, 503 or slow responses and growing by one on fast successes is the additive-increase, multiplicative-decrease rule, and it converges on what each host tolerates. Honour Retry-After when present, and combine the limit with the host's Crawl-delay from respecting robots.txt and crawl-delay asynchronously. The worker count for the host then follows limit — start or retire workers as it changes.
Verify: against a test host that returns 429 above three concurrent requests, the crawler settles at three without sustained errors.
Verification¶
Per-host limiting is correct when:
- No host ever sees more than its limit of concurrent requests.
- A slow host does not slow the others, because each host has its own queue.
- Active hosts and total sockets are capped, consistently with the connector.
- Limits adapt to 429s, 503s and latency for long crawls.
Diagnostic Hook: chart in-flight requests and fetch rate per host. In-flight piling up on one host under a global limit means the per-host structure is missing; workers blocked on semaphores while other queues are non-empty means a shared queue is causing head-of-line blocking.
Pitfalls & edge cases¶
- Only a global limit. Measured: the slow host received all 20 requests.
- Per-host semaphores with one shared queue. Measured: 16.28 s against 1.24 s.
- Truthiness checks on defaultdicts. An empty one is falsy and skips the limit.
- Unbounded hosts. Workers per host multiply with discovered hosts.
Frequently Asked Questions¶
How do I limit requests per host in an asyncio crawler?
Give each host its own queue and a fixed number of workers that only fetch from it, and set aiohttp's limit_per_host to the same number as a backstop. In testing that crawled in 1.24 s with every host capped at 4.
Why is my crawler slow with per-host semaphores?
With one shared queue, workers take a URL for a slow host and wait on its semaphore while URLs for other hosts wait behind them. In testing that took 16.28 s instead of 1.24 s with per-host queues.
What concurrency should a crawler use per site?
Start low, around 2 to 4, respect Crawl-delay, and adapt: halve on 429, 503 or slow responses and increase slowly on fast successes.
Is aiohttp's limit_per_host enough for a polite crawler?
It caps connections per host, but with a shared work queue it causes the same head-of-line blocking. Use it as a backstop behind per-host queues.
Related¶
- Async Web Crawlers — up to the topic overview.
- Deduplicating URLs in a concurrent crawl frontier — what goes into the per-host queues.
- Network I/O & Protocol Handling — the section overview.