Skip to content

Avoiding Crawl Traps and Infinite URL Spaces

Many sites generate more URLs than they have pages: calendars with a "next month" link forever, session IDs appended to every link, "next page" links on empty result pages, faceted filters in every combination, and relative links that nest one level deeper on each visit. An async crawler follows them all very efficiently. Measured against a local test site with 551 real pages — a home page, 50 section pages and 500 articles — plus all five of those traps, crawled by an aiohttp crawler with 20 workers and a budget of 5,000 fetches: with exact-URL deduplication only, it spent the whole budget, 4,500 fetches of it in traps — 1,456 calendar months, 1,451 copies of the home page with different session IDs, 1,474 facet and pagination pages — and would never have finished. Removing session parameters moved the waste rather than removing it: 4,142 facet and pagination fetches. Adding a cap of 100 fetches per URL pattern with query parameters let the crawl finish on its own after 951 fetches; adding a content fingerprint brought it to 902, with all 500 articles found. A fingerprint computed over too little of the page did the opposite: it judged every section page a duplicate and the crawl found 10 of the 500 articles. This guide builds these guards and shows where each one fails.

Prerequisites

1. Measure where the crawl budget goes

You cannot tell from the outside that a crawl is trapped — it is busy, fetching pages, finding links. Classify every fetch by URL shape and report the split, so trap traffic is visible:

def classify(url: str) -> str:
    path = urlsplit(url).path
    if re.fullmatch(r"/article/\d+/\d+", path):
        return "article"
    if "/related" in path:
        return "related loop"
    if path.startswith("/calendar"):
        return "calendar"
    if path.startswith("/section"):
        return "section/facets/pages" if urlsplit(url).query else "section"
    return "home"

On a real crawl the classes are not known in advance, so classify by pattern instead — the path with numbers replaced by N and the sorted names of the query parameters — and report the patterns with the most fetches. Measured with exact-URL deduplication only: 5,000 fetches, of which 500 were articles and 4,500 were traps; the crawl stopped only because the budget ran out. Every trap type contributed, and the article count looked fine — all 500 were found early — which is why trap traffic so often goes unnoticed until the crawl's cost does.

Verify: your crawler reports fetches per URL pattern, and you know what share of the budget went to patterns without useful content.

One 551-page site with five traps, 5,000-fetch budget A grid of 6 rows by 5 columns. One 551-page site with five traps, 5,000-fetch budget guards fetches finished? articles trap fetches exact-URL dedup 5,000 no, budget spent 500 4,500 + drop session params 5,000 no 500 4,500 (4,142 facets/pages) + path depth/repeat guard 5,000 no 500 4,500 + cap 100 per query pattern 951 yes 500 451 + content fingerprint (h1 + text) 902 yes 500 402 fingerprint over <p> text only 80 yes 10 70 aiohttp, 20 workers; traps: calendar, session IDs, empty pagination, facets, relative-link loop.

2. Normalise away parameters that do not change the page

Session IDs, tracking parameters and parameter order produce endless spellings of the same page. Strip what never changes content and sort the rest before deduplicating:

DROP_PARAMS = {"sid", "sessionid", "phpsessid", "utm_source", "utm_medium", "utm_campaign"}

def normalise(url: str) -> str:
    p = urlsplit(url)
    query = sorted((k, v) for k, v in parse_qsl(p.query) if k.lower() not in DROP_PARAMS)
    return urlunsplit((p.scheme, p.netloc.lower(), p.path or "/", urlencode(query), ""))

Measured: the 1,451 session-ID copies of the home page fell to 1 fetch, and calendar fetches fell from 1,456 to 15 — not because the calendar was solved, but because the budget now went to facets and pagination instead, at 4,142 fetches, and the crawl still did not finish. Normalisation is necessary, and on its own it does not bound anything: it removes duplicate spellings of finite pages, while traps create genuinely different URLs without end. The parameter list is per site; maintain it from the pattern report of step 1.

Verify: the same page reached with different session IDs or parameter orders is fetched once.

3. Reject paths that only a bug would produce

Relative links that resolve one level deeper each time — related/ on a page served at /article/1/2/related/ — produce paths like /article/1/2/related/related/related/. Real sites rarely nest deeply and almost never repeat a segment many times:

def path_ok(url: str, max_depth: int = 6, max_repeat: int = 2) -> bool:
    segments = [s for s in urlsplit(url).path.split("/") if s]
    return len(segments) <= max_depth and max(Counter(segments).values(), default=0) <= max_repeat

Measured with normalisation already on: the relative-link loop fell from 292 fetches to 100, at most two levels per article. The crawl still spent its budget, now on pagination and facets, which have short, valid-looking paths. Path guards are cheap and specific; set the limits from the deepest legitimate URLs in the site's sitemap, so they never reject real pages.

Verify: no fetched path exceeds the depth limit or repeats a segment more than allowed, and no sitemap URL would be rejected.

4. Cap fetches per URL pattern

The guard that made the crawl finish was a budget per URL pattern. Calendars, pagination and facets each produce unlimited URLs of one shape, so capping each shape bounds them all at once:

QUERY_CAP, PATH_CAP = 100, 5000

def pattern(url: str) -> str:
    p = urlsplit(url)
    shape = re.sub(r"\d+", "N", p.path)
    keys = ",".join(sorted(k for k, _ in parse_qsl(p.query)))
    return f"{shape}?{keys}" if keys else shape

def within_budget(url: str, counts: Counter) -> bool:
    pat = pattern(url)
    if counts[pat] >= (QUERY_CAP if "?" in pat else PATH_CAP):
        return False
    counts[pat] += 1
    return True

Measured: with a cap of 100 for patterns that have query parameters, the crawl ran out of URLs on its own after 951 fetches — 500 articles, 51 hubs, 100 calendar months, 200 facet and pagination pages, 100 loop pages — instead of exhausting 5,000. A first attempt with a cap of 1,000 finished too, but only after 3,251 fetches, 1,000 of them calendar months. Caps cut both ways: content that is reachable only through a capped pattern, such as articles listed only on paginated pages, is lost past the cap. That is why path-only patterns like /article/N/N get a much larger cap here, and why every pattern that hits its cap should be logged — a cap reached by a real-content pattern is a signal to raise it or to use the site's sitemap instead, as in discovering URLs from sitemaps asynchronously.

Verify: the crawl of a trap-heavy site finishes before its budget, and the list of patterns that hit their cap contains only trap patterns.

Which guard stops which trap? A decision on What generates the extra URLs with 4 outcomes. Which guard stops which trap? What generates the extra URLs? session / tracking parameters normalise: drop + sort 1,451 copies to 1 relative links nesting deeper depth + repeated-segment limit 292 to 100 calendar, facets, pagination cap per URL pattern finished at 951 same page, different URL content fingerprint 951 to 902 Only the per-pattern cap made the crawl finite on its own.

5. Stop expanding pages you have already seen, carefully

Empty pagination pages, facet combinations that show the same results and loop pages all return content the crawler has already seen. Fingerprint the page's text and do not follow links from a page whose fingerprint is known — it can only lead to more of the same:

def fingerprint(tree) -> bytes:
    text = " ".join(node.text() for node in tree.css("h1, p"))
    return hashlib.blake2b(text.encode(), digest_size=16).digest()

def should_expand(tree, seen_content: set[bytes]) -> bool:
    fp = fingerprint(tree)
    if fp in seen_content:
        return False                         # fetched and counted, but not expanded
    seen_content.add(fp)
    return True

Measured on top of the other guards: 902 fetches instead of 951, still with all 500 articles. The fingerprint must cover what distinguishes pages from each other. A first version hashed only paragraph text; every section page on the test site shared the same introductory paragraph, so after the first section the crawler judged the other 49 duplicates and never followed their article links — 80 fetches, 10 articles. Include headings and main content, exclude navigation and boilerplate, and check the fingerprint against pages you know are distinct. Near-duplicate detection such as SimHash catches pages that differ only by a timestamp or counter, at the cost of more tuning. Persist both the URL set and the pattern counts with the rest of the crawl state, as in persisting crawl state to resume after a crash, or a resumed crawl starts its traps afresh.

Verify: on a sample of known-distinct pages, no two fingerprints collide.

Guards on every discovered link and fetched page A flow of 5 stages. Guards on every discovered link and fetched page Normalise drop session params, sort query Path guard depth <= 6, repeats <= 2 Seen URL? exact match after normalising Pattern cap 100 per query pattern Fingerprint do not expand duplicates The first four run per link; the fingerprint runs per fetched page.

Verification

A crawler is protected from traps when:

  • Fetches are reported per URL pattern, and trap patterns are visible.
  • Session and tracking parameters are removed and query parameters sorted.
  • Path depth and repeated segments are limited, below any real URL's depth.
  • Each query pattern has a fetch cap, and caps reached are logged; content fingerprints cover distinguishing text.

Diagnostic Hook: alert on any host where one URL pattern accounts for more than half of the fetches, or where the number of distinct patterns keeps growing. Healthy crawls converge on a handful of patterns per site; a trap shows up as one pattern eating the budget or as patterns multiplying with every page.

Pitfalls & edge cases

  • Exact-URL deduplication as the only guard. Measured: 4,500 of 5,000 fetches in traps.
  • Normalisation alone. The waste moved to facets and pagination.
  • Caps that hit real content. Log every pattern at its cap.
  • Fingerprints over too little text. Measured: 10 of 500 articles found.

Frequently Asked Questions

What is a crawler trap?

A part of a site that generates unlimited distinct URLs, such as calendars, session IDs in links, empty pagination, faceted filters and relative-link loops. On a 551-page test site, they took 4,500 of a 5,000-fetch budget.

How do I stop a crawler following infinite calendars and pagination?

Cap fetches per URL pattern: the path with numbers replaced and the names of query parameters. A cap of 100 per query pattern made the test crawl finish after 951 fetches.

Is removing session IDs from URLs enough?

No. It removed 1,450 duplicate home pages, but the budget then went to facets and pagination instead, and the crawl still did not finish.

Can content hashing replace URL rules?

It helps, reducing 951 fetches to 902, but it only acts after a page is fetched, and a hash over too little text marked distinct pages as duplicates and lost 490 of 500 articles.