Skip to content

Parsing HTML Off the Event Loop

An async crawler spends most of its wall time waiting on the network and most of its CPU time parsing HTML, and the parsing runs on the event loop unless you move it. How to move it depends on the parser: some release the GIL and run in parallel in threads, some do not, and for some the cost of shipping a page to another process exceeds the parse itself. Measured on Python 3.14 with a 194 KB catalogue page containing 900 links, extracting every href: BeautifulSoup with html.parser took 52–75 ms per page and BeautifulSoup with lxml 41–46 ms; lxml's own API with XPath took 1.8–1.9 ms and selectolax's Lexbor parser 3.3 ms. Parsing on the loop, BeautifulSoup produced maximum loop lags of 123–263 ms. In a four-thread pool, selectolax rose from about 300 to 913–971 pages/s — it releases the GIL — while lxml fell from about 500 to 180–214 pages/s and BeautifulSoup gained nothing and still caused 195–227 ms of lag. A four-process pool took BeautifulSoup from 14–27 to 71–94 pages/s with lag around 1–2 ms. This guide shows how to measure your parser and place it.

Prerequisites

1. Measure what one page costs to parse

Start with the number that decides everything else: CPU time per page with your parser, on pages like the ones you crawl. Extract exactly what the crawler needs — usually links and a few fields — because building a full object tree costs more than reading attributes:

from bs4 import BeautifulSoup
import lxml.html
from selectolax.lexbor import LexborHTMLParser

def bs4_links(html):        return [a["href"] for a in BeautifulSoup(html, "html.parser").find_all("a", href=True)]
def lxml_links(html):       return lxml.html.fromstring(html).xpath("//a/@href")
def selectolax_links(html): return [n.attributes["href"] for n in LexborHTMLParser(html).css("a[href]")]

Measured as the median of 15 runs, in two sessions: html.parser through BeautifulSoup 52.2–75.4 ms, BeautifulSoup over lxml 41.2–45.8 ms, lxml XPath 1.8–1.9 ms, selectolax 3.3 ms; all four found the same 900 links. At 60 ms a page, one core parses about 16 pages a second — far less than an async fetcher can download — so with BeautifulSoup the parser, not the network, is the crawler's ceiling. The machine was shared with other work during these runs, which is why the slow parsers vary most.

Verify: you know your parser's per-page time on representative pages and how many pages per second one core can parse.

Milliseconds to extract 900 links from a 194 KB page 4 horizontal bars comparing BeautifulSoup, html.parser with the others. Milliseconds to extract 900 links from a 194 KB page BeautifulSoup, html.parser 52-75 ms BeautifulSoup, lxml 41-46 ms selectolax (Lexbor) 3.3 ms lxml.html + XPath 1.8-1.9 ms bs4 4.15.0, lxml 6.1.3, selectolax 0.4.13; Python 3.14; median of 15, two sessions. The parser choice spans a factor of 40.

2. Find out whether the parser releases the GIL

asyncio.to_thread only helps a CPU-bound call if the call releases the GIL while it works. Test it directly: parse the same pages on the loop, in a thread pool, and in a process pool, and measure throughput and loop lag in each case:

async def compare(fn, html, pages=48):
    loop = asyncio.get_running_loop()
    threads, procs = ThreadPoolExecutor(4), ProcessPoolExecutor(4)

    async def on_loop():
        for _ in range(pages):
            fn(html)
            await asyncio.sleep(0)

    async def in_threads():
        await asyncio.gather(*(loop.run_in_executor(threads, fn, html) for _ in range(pages)))

    async def in_procs():
        await asyncio.gather(*(loop.run_in_executor(procs, fn, html) for _ in range(pages)))
    ...                                   # time each with a loop-lag probe running alongside

Measured in two sessions: selectolax went from 297–303 pages/s on the loop to 913–971 in four threads, with loop lag 2.5–6.2 ms — Lexbor parses without holding the GIL. lxml went the other way, from 464–538 pages/s on the loop to 180–214 in threads, with lag rising to 27–44 ms: its parse holds the GIL, so four threads contended for one lock and the handoffs cost more than the 2 ms parse. BeautifulSoup is mostly Python code; four threads gave 9–15 pages/s, no better than the loop, and lag stayed at 37–227 ms because the worker threads kept the loop thread waiting for the GIL. The result is specific to each library and version, which is why it must be measured rather than assumed.

Verify: for your parser, threads either raise throughput and keep lag low, or you have ruled them out.

Pages per second and max loop lag, by where parsing runs A grid of 4 rows by 4 columns. Pages per second and max loop lag, by where parsing runs parser on the loop 4 threads 4 processes BeautifulSoup, html.parser 14-16/s, lag 221-263 ms 13-15/s, lag 195-227 ms 71/s, lag 1.6-2.1 ms BeautifulSoup, lxml 20-27/s, lag 123-159 ms 9-11/s, lag 37-53 ms 81-94/s, lag 1.0-1.1 ms lxml + XPath 464-538/s, lag 6-8 ms 180-214/s, lag 27-44 ms 473-512/s, lag 17-22 ms selectolax 297-303/s, lag 8-10 ms 913-971/s, lag 2.5-6.2 ms 824-851/s, lag 1.0 ms Two sessions of 48 pages each; ranges show both.

3. Pick a parser for the crawler, not for the tutorial

BeautifulSoup's API is pleasant and its tolerance for broken markup is good, but at 40–75 ms a page it turns a crawler that could fetch hundreds of pages a second into one that processes a few dozen. For link extraction and simple field scraping, a fast parser with CSS selectors or XPath does the same job:

from selectolax.lexbor import LexborHTMLParser

def extract(html: str, base_url: str) -> tuple[list[str], dict]:
    tree = LexborHTMLParser(html)
    links = [n.attributes["href"] for n in tree.css("a[href]")]
    title = tree.css_first("title")
    return links, {"title": title.text(strip=True) if title else None}

Return plain data — strings, lists, dicts — rather than parser nodes, so the result can cross a thread or process boundary and the parse tree can be freed immediately. Resolve relative links with urllib.parse.urljoin(base_url, href) and normalise them before they reach the frontier, as in deduplicating URLs in a concurrent crawl frontier. Keep BeautifulSoup for the pages that genuinely need its repairs, and move those to a process pool.

Verify: the fast parser extracts the same links and fields as your previous one on a sample of real pages.

4. Use a process pool when the parser holds the GIL

For parsers that hold the GIL and cost tens of milliseconds, processes are the only way to parse in parallel without stalling the loop. Send the page and return only the extracted data:

PARSE_POOL = ProcessPoolExecutor(max_workers=4)        # created once, at start-up

async def parse_page(html: str, url: str):
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(PARSE_POOL, extract_with_bs4, html, url)

Measured: BeautifulSoup in four processes reached 71–94 pages/s with loop lag of 1.0–2.1 ms, against 14–27 pages/s and 123–263 ms on the loop. For lxml, processes gave 473–512 pages/s — no better than parsing on the loop — with 17–22 ms of lag: pickling a 194 KB page and 900 strings each way cost as much as the 2 ms parse it was meant to offload. A process pool pays off when the parse costs much more than serialising its input and output; for a fast parser that holds the GIL, parsing on the loop, one page at a time between awaits, can be the best option. Reduce what crosses the boundary by sending bytes rather than decoded text and returning only what the crawler uses, as covered in reducing pickle overhead in ProcessPoolExecutor payloads.

Verify: throughput with the pool exceeds throughput on the loop, and loop lag during crawling stays within budget.

Where should this parser run? A decision on What does the parser do with the GIL with 4 outcomes. Where should this parser run? What does the parser do with the GIL? releases it (selectolax) thread pool 913-971 pages/s holds it, ~2 ms/page (lxml) on the loop, between awaits 464-538 pages/s holds it, ~50 ms/page (BeautifulSoup) process pool 71-94 pages/s, lag 1-2 ms parse cheaper than pickling not a process pool lxml: no gain Measure throughput and loop lag; neither alone is enough.

5. Size the parse stage against the fetch stage

In the crawler, parsing is a pipeline stage between fetching and the frontier. Give it a bounded number of workers equal to the executor's size, so pages queue in a bounded queue instead of as thousands of pending futures:

async def parse_stage(pages: asyncio.Queue, frontier, pool, workers=4):
    loop = asyncio.get_running_loop()

    async def worker():
        while (item := await pages.get()) is not None:
            url, html = item
            links, fields = await loop.run_in_executor(pool, extract, html, url)
            await frontier.add_many(links)
            await store(url, fields)

    async with asyncio.TaskGroup() as tg:
        for _ in range(workers):
            tg.create_task(worker())

When parsing is the slower stage, the bounded page queue fills and fetchers wait — which is correct: downloading pages faster than they can be parsed only buys memory. Compare the stages' rates as in scaling a slow pipeline stage: with selectolax, four threads parsed about 950 pages/s, more than most polite crawlers fetch; with BeautifulSoup, four processes parsed under 100, which may well be the crawler's limit.

Verify: the page queue between fetch and parse stays bounded, and loop lag stays low while it is full.

Verification

HTML parsing is off the loop correctly when:

  • Per-page parse time is known for your parser on representative pages.
  • The executor matches the parser: threads only if it releases the GIL, processes only if the parse outweighs pickling.
  • Workers return plain data, never parser objects.
  • The parse stage is bounded, with a queue between it and the fetchers.

Diagnostic Hook: record parse time per page alongside fetch time, and alert on the slowest parses. Pathological pages — megabytes of markup, deeply nested tables — show up as parse outliers of hundreds of milliseconds; cap the page size you accept, and send oversized pages to a separate pool so they cannot delay the rest.

Pitfalls & edge cases

  • Parsing with BeautifulSoup on the loop. Measured: 123–263 ms of loop lag.
  • Assuming threads help. lxml fell from about 500 to about 200 pages/s in threads.
  • Process pools for fast parsers. Pickling cost as much as the parse.
  • Returning parse trees from workers. They cannot cross processes and keep memory alive.

Frequently Asked Questions

Should I parse HTML in asyncio.to_thread?

Only if the parser releases the GIL. selectolax did, reaching 913 to 971 pages/s in four threads; lxml and BeautifulSoup did not, and lxml got slower in threads.

What is the fastest HTML parser for an async crawler?

In testing, lxml with XPath (1.8 to 1.9 ms per page) and selectolax (3.3 ms) against 41 to 75 ms for BeautifulSoup. selectolax also scaled across threads.

Does BeautifulSoup block the event loop?

Yes: each page took 41 to 75 ms of CPU, and parsing on the loop produced 123 to 263 ms of loop lag. A four-process pool kept lag at 1 to 2 ms and parsed 71 to 94 pages/s.

When is a process pool not worth it for parsing?

When the parse is cheaper than pickling the page and the result: lxml in processes was no faster than on the loop.