Skip to content

Adapting Crawl Rate to Host Latency

A crawler that uses one concurrency limit for every host is wrong for nearly all of them: polite enough for the smallest site means needlessly slow on a large one, and fast enough for the large one means overloading the small. The host's response times already say how much load it can take. Measured on Python 3.14 with aiohttp against two simulated hosts — host A with 8 parallel workers at 20 ms per page (capacity 400 pages/s), host B with 2 workers at 50 ms (40 pages/s), each shedding load with 503 Retry-After: 1 when its queue was full: a fixed limit of 2 crawled A at 95 pages/s — under a quarter of its capacity. A fixed limit of 32 crawled A at 283 pages/s with 256 503 responses and B at just 19 pages/s with 658 503s, slower than being polite. An adaptive limit that grew while latency stayed near the host's best and shrank on slowdowns and 503s crawled A at 365–371 pages/s and B at 39 pages/s — both near capacity — with zero 503s, settling at limits of about 9–12 and 3–5. When host A slowed to 80 ms per page mid-crawl, an adaptive limit that remembered the host's best latency forever collapsed to 1 and crawled at 12 pages/s; one whose baseline expired after a window held 82–86 pages/s against a new capacity of 100. This guide builds that limiter.

Prerequisites

1. See what fixed limits cost

Before adapting, measure the two failure modes of a fixed limit. Too low leaves capacity unused; too high fills the host's queue, which raises latency for everyone, and once the host starts shedding load, the retries make the crawl slower than if it had been polite:

class Fixed:
    def __init__(self, n):
        self.sem = asyncio.Semaphore(n)
    async def __aenter__(self):
        await self.sem.acquire()
    async def __aexit__(self, *exc):
        self.sem.release()

Measured with 2,000 pages on host A and 300 on host B: a limit of 2 gave 95 and 39 pages/s with no queueing — perfect for B, a quarter of A's capacity. A limit of 32 pushed host A's queue to its 16-request maximum, raised median latency from 21 to 63 ms, and drew 256 503s while reaching 283 pages/s; on host B it drew 658 503s, each followed by a one-second wait, and finished at 19 pages/s — half of what the limit of 2 achieved. No single number serves both hosts, and real crawls touch thousands.

Verify: for a sample of hosts, you know the throughput and error rate of your current fixed limit.

Two hosts, three limits A grid of 4 rows by 5 columns. Two hosts, three limits limit host A (cap. 400/s) host B (cap. 40/s) 503s settled limit A / B fixed 2 95 pages/s, p50 21 ms 39 pages/s, p50 51 ms 0 2 / 2 fixed 32 283 pages/s, p50 63 ms 19 pages/s, p50 153 ms 256 + 658 32 / 32 adaptive, tolerance 2.0 341-369 pages/s 39 pages/s, p50 101 ms 0 10.5-11.5 / 3.8-4.7 adaptive, tolerance 1.5 371 pages/s, p50 21 ms 39 pages/s, p50 76 ms 0 9.0 / 3.2 aiohttp client, simulated hosts shedding load with 503 Retry-After: 1.

2. Grow while latency holds, shrink when it rises

A host's latency at low load is its service time; when requests start queueing, latency rises before errors appear. Use that as the signal: keep a per-host limit, raise it slowly while latency stays within a tolerance of the host's best, lower it when latency exceeds it, and halve it on an explicit overload response:

class AdaptiveLimit:
    def __init__(self, initial=2, minimum=1, maximum=64, tolerance=2.0):
        self.limit, self.minimum, self.maximum, self.tolerance = float(initial), minimum, maximum, tolerance
        self.in_flight = 0
        self.best: float | None = None
        self.paused_until = 0.0
        self._cond = asyncio.Condition()

    async def __aenter__(self):
        async with self._cond:
            await self._cond.wait_for(lambda: self.in_flight < int(self.limit))
            self.in_flight += 1
        if (delay := self.paused_until - time.monotonic()) > 0:
            await asyncio.sleep(delay)

    async def __aexit__(self, *exc):
        async with self._cond:
            self.in_flight -= 1
            self._cond.notify_all()

    def record(self, latency: float, status: int, retry_after: float | None = None):
        if status in (429, 503):
            self.limit = max(self.minimum, self.limit / 2)
            self.paused_until = time.monotonic() + (retry_after or 1.0)
            return
        self.best = latency if self.best is None else min(self.best, latency)
        if latency > self.best * self.tolerance:
            self.limit = max(self.minimum, self.limit * 0.9)
        else:
            self.limit = min(self.maximum, self.limit + 1 / self.limit)

Adding 1 / limit per response raises the limit by about one per round of requests — the additive-increase, multiplicative-decrease shape that TCP congestion control uses. Measured: host A settled at a limit of about 10–11 and 341–369 pages/s with no 503s; host B at about 4 and 39 pages/s, its capacity.

Verify: on a host with known capacity, the limit settles near capacity × service time and 503s stay at zero.

3. Choose the tolerance as a politeness setting

The tolerance is how much extra latency you are willing to cause. At 2.0, the crawler accepts doubling a host's response time: measured on host B, median latency rose from 51 ms to 101 ms with the limit around 4, for the same 39 pages/s that a limit of 2 achieved at 51 ms. That extra latency is queueing inside the host — felt by its other users too — for no gain once the host is at capacity. At 1.5, host B settled at 3.2 with a 76 ms median and host A at 9.0 with a 21 ms median and 371 pages/s:

limiter = AdaptiveLimit(tolerance=1.5)       # accept at most ~50% added latency

Lower tolerances leave more headroom for the host's real users and react sooner to load from elsewhere; higher ones squeeze out a little more throughput from hosts with spare parallelism. For a general-purpose crawler, a tolerance between 1.2 and 1.5 errs on the side of the host. The limit also needs a ceiling: maximum stops a fast, idle host from absorbing hundreds of connections, which would break the per-host fairness covered in limiting concurrency per host in an async crawler.

Verify: at the settled limit, median latency is within your tolerance of the host's unloaded latency.

Host A slows down mid-crawl 3 lanes over time. Host A slows down mid-crawl host A service time 20 ms 80 ms 20 ms best kept forever ~350/s 12/s ~370/s best over 1 s window ~375/s 82-86/s ~370/s time → A baseline that never expires treats a new normal as permanent overload.

4. Let the baseline expire

The limiter compares latency against the host's best, and "best" has to mean recent best. Hosts change speed — a cache warms, a backup starts, traffic shifts — and a baseline from an hour ago turns a permanent slowdown into a permanent emergency. Measured when host A's service time rose from 20 to 80 ms between seconds 3 and 6: with the best latency kept forever, every response looked four times too slow, the limit fell to 1 and stayed there, and the crawl ran at 12 pages/s for the whole slowdown. Keep the minimum over a sliding window instead:

def record_latency(self, latency: float) -> None:
    now = time.monotonic()
    if now - self._window_start > self.window:            # e.g. 1-30 s
        self._previous_min, self._current_min = self._current_min, None
        self._window_start = now
    self._current_min = latency if self._current_min is None else min(self._current_min, latency)
    self.best = min(x for x in (self._current_min, self._previous_min) if x is not None)

With a one-second window the limit dropped briefly, then recovered to about 10 once the 80 ms latency became the baseline, and the crawl held 82–86 pages/s against the host's reduced capacity of 100 — then returned to about 370 when the host recovered. For comparison, a fixed limit of 10, which happens to suit host A, ran at 96–100 pages/s during the slowdown — but the same fixed limit would overload host B four times over. In real crawls, use a window of tens of seconds; it should be long enough to contain some unloaded samples and short enough to forget an old normal.

Verify: after a host's latency changes permanently, the limit recovers to a new steady value within a few windows.

5. Honour explicit overload signals first

Latency is an inference; 429 and 503 with Retry-After are the host telling you directly. Treat them as stronger than any latency signal: halve the limit, pause the host for the time it asked for, and retry the page later rather than immediately:

async def fetch_page(session, url, limiter):
    while True:
        async with limiter:
            start = time.perf_counter()
            async with session.get(url) as resp:
                body = await resp.read()
                retry_after = parse_retry_after(resp.headers.get("Retry-After"))
                limiter.record(time.perf_counter() - start, resp.status, retry_after)
        if resp.status not in (429, 503):
            return resp.status, body
        await asyncio.sleep(retry_after or 1.0)       # wait outside the limiter

The fixed limit of 32 ignored these signals except for the per-request wait, and spent most of its time in those waits: 658 503s on a 300-page crawl. The adaptive limiter never drew one, because latency rose and the limit came down before the host's queue filled. Parse Retry-After in both its forms — seconds and an HTTP date — as covered in handling 429 Retry-After responses in async clients, and combine the adaptive limit with any Crawl-delay from robots.txt, taking whichever is more conservative.

Verify: a host that returns 503 causes an immediate halving of its limit and a pause of at least Retry-After.

What should the limiter do with this response? A decision on What came back with 4 outcomes. What should the limiter do with this response? What came back? 429 / 503 halve, pause for Retry-After 0 such responses when adaptive slow: > tolerance x recent best limit x 0.9 host is queueing normal latency limit + 1/limit about +1 per round any response cap at maximum and Crawl-delay politeness first Explicit signals override inferred ones.

Verification

A crawler adapts to host latency when:

  • Each host has its own limit, starting low and growing while latency holds.
  • Latency is compared with a recent best, over a window that expires.
  • 429 and 503 halve the limit and pause the host for Retry-After.
  • Tolerance and maximum are set as politeness choices, not left at defaults.

Diagnostic Hook: export each host's current limit and recent best latency. A host whose limit sits at the minimum for long periods is either overloaded or has a baseline that no longer fits; a host pinned at the maximum has capacity you have chosen not to use. Both are worth a look, and neither shows up in aggregate pages per second.

Pitfalls & edge cases

  • One fixed limit for all hosts. Measured: a quarter of capacity on one host, 658 503s on another.
  • A best latency that never expires. Measured: 12 pages/s through a slowdown that allowed 100.
  • High tolerance. Doubled host latency for no extra throughput.
  • Retrying a 503 inside the limiter. The wait occupies a slot; sleep outside it.

Frequently Asked Questions

How do I choose crawl concurrency per host?

Let it adapt: start low, add about one concurrent request per round while latency stays near the host's recent best, and back off on slowdowns and 503s. In testing this settled near each host's capacity (about 10 and 4) with no 503s.

Why is a high fixed concurrency slower on small sites?

The host's queue fills, it sheds load with 503s, and every retry waits. A limit of 32 crawled a small host at 19 pages/s with 658 503s, half the speed of a limit of 2.

What latency increase should an adaptive crawler tolerate?

Treat it as politeness: a tolerance of 1.5 kept median latency at 21 ms on a large host and 76 ms on a small one with full throughput; 2.0 doubled the small host's latency for no gain.

What should a crawler do on 503 Retry-After?

Halve that host's limit, pause the host for at least the Retry-After time, and retry the page later outside the limiter.