Discovering URLs from Sitemaps Asynchronously¶
A site's sitemaps list its pages directly, which saves a crawler from discovering them link by link — but a large site's sitemaps are an index of gzipped files holding up to 50,000 URLs each, and fetching, decompressing and parsing them is real work. Measured on Python 3.14 against a local server with 50 ms of latency per request, serving a sitemap index of 20 gzipped sitemaps — 1,000,000 URLs, 4.1 MB of XML per file: fetching and parsing them one after another with ElementTree on the loop took 4.46–5.06 s with loop lag up to 227–305 ms. Fetching four at a time and parsing with asyncio.to_thread was only slightly faster, 2.90–3.72 s, and made lag worse — 460–711 ms — because the parsing threads competed with the loop for the GIL. Parsing in a four-process pool took 1.15–1.24 s with 52–64 ms of lag. Filtering entries by lastmod inside the workers, for an incremental crawl, returned 50,095 recently changed URLs in 0.77 s with 18 ms of lag and a 59 MiB peak. A 194 KB gzip bomb expanding to 200 MB peaked the naive reader at 823 MiB and stalled the loop 1.6 s; a size-limited decompressor rejected it in 0.21 s. This guide builds the fast, safe version.
Prerequisites¶
- Python 3.11+, aiohttp or httpx, and a
ProcessPoolExecutor. - robots.txt handling, from respecting robots.txt and Crawl-delay asynchronously.
- The topic overview, Async Web Crawlers.
1. Find sitemaps through robots.txt¶
Sitemaps are announced with Sitemap: lines in robots.txt, which the crawler already fetches once per host. Read them from there rather than guessing /sitemap.xml:
def sitemap_urls(robots_txt: str) -> list[str]:
return [line.split(":", 1)[1].strip()
for line in robots_txt.splitlines()
if line.lower().startswith("sitemap:")]
A Sitemap: line applies to the whole file, not to a user-agent group, and a host may list several. Each URL can point to a <urlset> of pages or a <sitemapindex> of further sitemaps, possibly gzipped — the reader must handle all four combinations. The test server's robots.txt pointed to one index of 20 gzipped sitemaps, the layout most large sites use, because the protocol limits a single sitemap to 50,000 URLs and 50 MB uncompressed.
Verify: for each host, sitemap URLs come from robots.txt, with /sitemap.xml only as a fallback.
2. Fetch child sitemaps concurrently, with depth and loop guards¶
An index's children are independent, so fetch them concurrently — with a small per-host limit, because a crawler that respects a host's capacity for pages should respect it for sitemaps too:
async def discover(session, index_url, executor, since=None, limit=4, max_depth=3):
sem = asyncio.Semaphore(limit)
seen: set[str] = set()
urls: list[tuple[str, str | None]] = []
loop = asyncio.get_running_loop()
async def visit(url: str, depth: int) -> None:
if url in seen or depth > max_depth:
return # loops and absurd nesting stop here
seen.add(url)
async with sem:
async with session.get(url) as resp:
resp.raise_for_status()
body = await resp.read()
kind, entries = await loop.run_in_executor(executor, parse_sitemap, body, since)
if kind == "sitemapindex":
async with asyncio.TaskGroup() as tg:
for loc, _ in entries:
tg.create_task(visit(loc, depth + 1))
else:
urls.extend(entries)
await visit(index_url, 0)
return urls
The seen set and depth limit are not optional. The test server also served an index that listed itself; a reader without them was still following it after 50 fetches, while the guarded reader fetched it once. Measured fetching four at a time: the network part of the job dropped from 20 sequential round trips to five rounds — but total time barely moved until parsing was moved too, which is the next step.
Verify: a self-referencing index is fetched once, and nesting beyond max_depth is ignored.
3. Parse sitemaps in processes, not threads¶
Each 50,000-URL sitemap is 4.1 MB of XML, and parsing it is CPU work that holds the GIL. Compared on one file: ElementTree's fromstring took 91.5 ms, ElementTree's iterparse 186.8 ms, lxml's fromstring with XPath 236.4 ms and lxml's iterparse 180–271 ms. The full parse was the fastest, and the protocol's 50 MB cap keeps a full tree affordable, so there is no need to stream. Clearing elements while streaming can even be a trap: one common pattern — deleting each element's earlier siblings as you go — slowed the whole job to 31 s.
NS = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
def parse_sitemap(body: bytes, since: str | None = None):
raw = gunzip_limited(body) if body[:2] == b"\x1f\x8b" else body
root = ET.fromstring(raw)
kind = root.tag.removeprefix(NS)
tag = "sitemap" if kind == "sitemapindex" else "url"
entries = []
for el in root.iter(NS + tag):
lastmod = el.findtext(NS + "lastmod")
if since is None or kind == "sitemapindex" or lastmod is None or lastmod >= since:
entries.append((el.findtext(NS + "loc"), lastmod))
return kind, entries
Run in threads, this made things worse: with four parses in flight the loop lagged up to 711 ms, because four threads holding the GIL starve the loop thread; one parse thread gave 171 ms. In a four-process pool, created once at start-up, the job took 1.15–1.24 s with 52–64 ms of lag — the remaining lag is unpickling 50,000 tuples per file on the loop, which the next step shrinks. Python's bundled expat (2.7.4 here) also rejected a billion-laughs entity payload with "limit on input amplification factor (from DTD and entities) breached", so ElementTree is safe against that attack on current versions.
Verify: loop lag during sitemap discovery stays within budget while throughput beats parsing on the loop.
4. Filter by lastmod inside the worker for incremental crawls¶
After the first full crawl, most URLs have not changed. Pass the cut-off into the worker and filter there, so unchanged entries never cross the process boundary or enter the frontier:
since = last_crawl_started.date().isoformat() # e.g. "2025-09-01"
urls = await discover(session, index_url, PARSE_POOL, since=since)
Measured with a cut-off that matched about 5% of entries: 50,095 URLs in 0.77 s, maximum loop lag 18 ms and peak RSS 59 MiB, against 1.15–1.24 s, 52–64 ms and 266 MiB for the full list. Comparing ISO dates as strings works when they share a format; normalise mixed date-time formats with datetime.fromisoformat in the worker. lastmod is advisory — sites set it inconsistently, and some never update it — so pair incremental sitemap reads with a periodic full crawl. Index-level lastmod values can also skip whole child sitemaps that have not changed since the last run, saving their fetches entirely.
Verify: an incremental run returns only entries modified since the cut-off, and a periodic full run catches pages whose lastmod was not updated.
5. Refuse oversized and malicious sitemaps¶
A sitemap is input from a server you do not control. A gzipped file can expand by a factor of a thousand, and gzip.decompress will allocate whatever it expands to:
import zlib
MAX_UNCOMPRESSED = 50 * 1024 * 1024 # the protocol's own limit
class TooLarge(Exception):
pass
def gunzip_limited(data: bytes, limit: int = MAX_UNCOMPRESSED) -> bytes:
d = zlib.decompressobj(wbits=31) # 31 = expect a gzip header
out = d.decompress(data, limit + 1)
if len(out) > limit or d.unconsumed_tail:
raise TooLarge(f"sitemap expands past {limit} bytes")
return out
Measured with a 194 KB gzip file that expanded to 200 MB: the naive reader decompressed and parsed it, peaking at 823 MiB and stalling the loop for 1.64 s; the limited reader raised TooLarge after 0.21 s at 145 MiB. Also cap the compressed download itself — read the body with a size limit, as in limiting response size in async clients — and give each sitemap fetch a timeout, so a slow or endless response cannot hold a slot.
Verify: a gzip bomb is rejected before it is fully decompressed, and peak memory stays near the 50 MB limit.
Verification¶
Sitemap discovery is fast and safe when:
- Sitemaps come from robots.txt, and indexes are followed with a seen set and depth limit.
- Child sitemaps are fetched concurrently within a per-host limit.
- Parsing runs in a long-lived process pool, with
lastmodfiltering inside the worker. - Decompression and downloads are size-limited, and each fetch has a timeout.
Diagnostic Hook: log, per host and run, the number of sitemap files, URLs found, URLs after the lastmod filter, and parse time. A host whose URL count jumps by orders of magnitude has started generating sitemap entries for an unbounded URL space — the same problem as a crawl trap, arriving through the sitemap.
Pitfalls & edge cases¶
- Parsing sitemaps in
to_thread. Measured: loop lag up to 711 ms, worse than the loop itself. gzip.decompresson untrusted input. Measured: 823 MiB for a 194 KB file.- Following indexes without a seen set. A self-referencing index never ends.
- Trusting
lastmodcompletely. Pair incremental runs with periodic full runs.
Frequently Asked Questions¶
How do I read sitemaps with asyncio?
Get sitemap URLs from robots.txt, fetch the index and its children concurrently with a small limit, and parse them in a process pool. 1,000,000 URLs across 20 gzipped files took 1.2 s this way in testing, against 4.5 to 5 s sequentially.
Should I parse sitemap XML in a thread?
In testing, four parse threads raised loop lag to 460 to 711 ms because they competed for the GIL; a process pool kept it at 52 to 64 ms and was faster.
How do I crawl only changed pages from a sitemap?
Filter entries by lastmod against your last crawl's start time inside the parsing worker. In testing, 50,095 of 1,000,000 URLs came back in 0.77 s. Run a full pass periodically, since lastmod is not always maintained.
How do I protect a crawler from gzip bombs in sitemaps?
Decompress with zlib.decompressobj and a maximum output size (the protocol allows 50 MB). A 194 KB bomb expanding to 200 MB was rejected in 0.21 s.
Related¶
- Async Web Crawlers — up to the topic overview.
- Parsing HTML off the event loop — the same placement question for pages.
- Network I/O & Protocol Handling — the section overview.