Skip to content

Pooling Browser Contexts for Concurrent Pages

A headless browser can render many pages at once, and an asyncio program can drive them with one task per page. The question is how many: too few and the browser waits on network round trips one page at a time; too many and memory climbs while throughput stops improving, because the CPU, the target site or Chromium's own scheduling becomes the bottleneck. Measured on Python 3.14 with Playwright 1.63 and headless Chromium, loading 100 JavaScript-rendered product pages from a local test site with a fresh context per page: 1 worker processed 2.3 pages per second; 4 workers, 8.0; 8 workers, 12.9; 16 workers, 12.7; 32 workers, 12.1. Peak memory (PSS across Chromium's processes) went 240 → 302 → 375 → 558 → 882 MiB. Past eight workers, every extra worker added about 20 MiB and no throughput. The machine was shared with other workloads during these runs, so absolute rates are approximate; the shape — a plateau with linearly rising memory — is the point. This guide finds the plateau and stays at it.

Prerequisites

1. Run a fixed pool of workers over a queue

Bound concurrency with a fixed number of worker tasks pulling URLs from a queue; each URL gets its own context:

import asyncio
from playwright.async_api import Browser


async def crawl(browser: Browser, urls: list[str], workers: int) -> dict[str, str]:
    queue: asyncio.Queue[str] = asyncio.Queue()
    for url in urls:
        queue.put_nowait(url)
    results: dict[str, str] = {}

    async def worker() -> None:
        while not queue.empty():
            url = queue.get_nowait()
            try:
                async with await browser.new_context() as context:
                    page = await context.new_page()
                    await page.goto(url)
                    await page.locator("#price[data-ready]").wait_for()
                    results[url] = await page.locator("#price").text_content()
            except Exception as exc:                        # one bad page must not stop a worker
                results[url] = f"error: {type(exc).__name__}"

    async with asyncio.TaskGroup() as tg:
        for _ in range(workers):
            tg.create_task(worker())
    return results

A fixed pool, rather than one task per URL with a semaphore, keeps the number of live contexts exactly equal to the number of workers, which is what memory depends on. Catching exceptions per URL keeps a timeout or a crashed page from cancelling the whole TaskGroup. The pool pattern is the same as in Worker Pool Implementations; the browser just makes each item expensive.

Verify: during a run, len(browser.contexts) never exceeds the worker count.

100 pages, a context per page, by worker count A grid of 5 rows by 4 columns. 100 pages, a context per page, by worker count workers pages/s peak Chromium PSS 100 pages took 1 2.3 240 MiB 44.2 s 4 8.0 302 MiB 12.5 s 8 12.9 375 MiB 7.7 s 16 12.7 558 MiB 7.9 s 32 12.1 882 MiB 8.2 s Playwright 1.63, Python 3.14; memory sampled every 0.5 s on a shared machine.

2. Find the plateau for your workload

The right worker count depends on what limits throughput — the target site, network latency, the page's own JavaScript, or the CPU running Chromium — so measure it rather than guessing:

async def find_plateau(browser: Browser, sample_urls: list[str]) -> None:
    for workers in (1, 2, 4, 8, 16, 32):
        start = time.perf_counter()
        await crawl(browser, sample_urls, workers)
        rate = len(sample_urls) / (time.perf_counter() - start)
        print(f"{workers:3d} workers: {rate:5.1f} pages/s, PSS {chrome_pss_mib():.0f} MiB")

In the measurement, pages per second rose almost linearly to four workers, reached its maximum at eight, and was flat or slightly lower beyond — while memory kept growing by about 20 MiB per worker. Pick the smallest worker count that reaches the plateau, then leave headroom below the memory limit for spikes: heavy pages can use several times the average. Sampling memory from a separate task, not inside the workers, matters for the measurement itself — an earlier version of this test that measured memory after every page slowed throughput noticeably, because walking Chromium's process tree takes tens of milliseconds.

Verify: your production worker count sits at the start of the plateau, and peak memory at that count fits the container limit with room to spare.

3. Prefer a context per page over long-lived pages

Reusing one page per worker, navigating it from URL to URL, skips context and page creation and seems cheaper. Measured over 400 pages with eight workers, it was not:

# Reused page per worker: no per-page setup, but state and memory accumulate
async def worker_reusing(context, queue):
    page = await context.new_page()
    while not queue.empty():
        await page.goto(queue.get_nowait())
        ...

A fresh context per page ran at 12.0 pages per second with PSS between 196 and 322 MiB throughout; one reused page per worker ran at 7.6 pages per second and its PSS climbed to 906 MiB by the 400th page. Long-lived pages accumulate history, cached resources and JavaScript heaps across navigations; a context per page throws all of that away every time. The isolation benefit — no cookies or storage carried between jobs — comes free. The memory side is covered in keeping headless browser memory bounded.

Verify: over a long run, Chromium's PSS stays in a band rather than trending upward.

400 pages with 8 workers, two lifecycles 2 horizontal bars comparing context per page: peak PSS with the others. 400 pages with 8 workers, two lifecycles context per page: peak PSS 322 MiB, 12.0 pages/s one reused page per worker: peak PSS 906 MiB, 7.6 pages/s Same machine and site; PSS sampled after every 100 pages. Fresh contexts were both lighter and faster than long-lived pages.

4. Rate-limit per host, not just overall

A worker pool bounds the browser's load; it does not bound the load on any one site. When crawling many hosts, eight workers can all land on the same host. Add a per-host limit in front of the browser:

from collections import defaultdict
from urllib.parse import urlsplit

PER_HOST = defaultdict(lambda: asyncio.Semaphore(2))


async def visit(browser: Browser, url: str) -> str:
    async with PER_HOST[urlsplit(url).hostname]:
        async with await browser.new_context() as context:
            page = await context.new_page()
            await page.goto(url)
            ...

Remember that one page load is many requests: the test pages fetched 22 resources each, so two concurrent pages on one host are about forty requests in flight against it. The per-host scheduling patterns from crawling apply unchanged, as in limiting concurrency per host in an async crawler, and blocking resources you do not need — measured in blocking and mocking requests in Playwright — reduces the load further.

Verify: the target host's request rate from your crawler stays within its documented or agreed limit.

5. Scale out with processes when one browser plateaus

A single Chromium instance shares its browser process across all contexts; once that or the event loop driving it saturates, more contexts do not help. Scaling past the plateau means more browsers in more processes:

from concurrent.futures import ProcessPoolExecutor


def run_shard(urls: list[str], workers: int) -> dict[str, str]:
    async def main() -> dict[str, str]:
        async with async_playwright() as p:
            browser = await p.chromium.launch()
            try:
                return await crawl(browser, urls, workers)
            finally:
                await browser.close()
    return asyncio.run(main())


with ProcessPoolExecutor(max_workers=4) as pool:
    shards = [urls[i::4] for i in range(4)]
    results = {}
    for part in pool.map(run_shard, shards, [8] * 4):
        results.update(part)

Each process runs its own event loop, Playwright driver and browser at its own plateau, and memory scales with processes × (browser base + workers × per-context cost): about 4 × (174 + 8 × 20) MiB here. This is the same shape as combining asyncio with multiprocessing for mixed workloads. Measure the plateau again with several processes, because the target site or the machine's CPU may become the limit first.

Verify: total throughput rises with the number of processes until CPU or the target, not a single browser, is the bottleneck.

How many browser workers should this crawl use? A decision on What does the measurement show with 4 outcomes. How many browser workers should this crawl use? What does the measurement show? throughput still rising add workers 2.3 -> 12.9 pages/s throughput flat, memory rising stop at the plateau 8 here many workers on one host per-host semaphore ~22 requests per page one browser saturated, CPU idle more processes, a browser each scale out The plateau, not the memory limit, sets the worker count.

Verification

Context pooling is sized well when:

  • A fixed pool of workers runs one context per page, and open contexts never exceed the pool size.
  • The worker count sits at the start of the measured throughput plateau.
  • Memory stays in a band over long runs, with fresh contexts rather than reused pages.
  • Per-host limits protect target sites, and extra capacity comes from processes.

Diagnostic Hook: plot pages per second and Chromium PSS against worker count from a short calibration run before each large crawl. The knee of the throughput curve is the worker count to use; if the knee moves between runs, the bottleneck is outside the browser — usually the target site — and per-host limits matter more than the pool size.

Pitfalls & edge cases

  • More workers past the plateau. Measured: 32 workers were no faster than 8 and used 882 MiB instead of 375.
  • Long-lived reused pages. Measured: 906 MiB after 400 pages and slower throughput.
  • Measuring memory inside workers. It slows the crawl it is measuring.
  • Overall limits only. Several workers can hit one host at once.

Frequently Asked Questions

How many concurrent pages can Playwright handle?

It depends on the workload; measure it. In testing, one headless Chromium reached its throughput plateau at about 8 concurrent contexts (12.9 pages/s); 16 and 32 were no faster and used up to 882 MiB of memory.

Should I reuse Playwright pages or create a new context per URL?

A new context per URL: over 400 pages it ran at 12.0 pages/s with memory in a 196–322 MiB band, while reusing one page per worker ran at 7.6 pages/s and grew to 906 MiB.

How do I limit concurrency in Playwright with asyncio?

Run a fixed number of worker tasks pulling URLs from an asyncio.Queue, each opening and closing its own context per URL, so live contexts never exceed the worker count.

How do I scale Playwright beyond one browser?

Run several processes, each with its own event loop, Playwright driver and browser, and split the URLs between them; each process runs at its own plateau.