Skip to content

Async Browser Automation with Playwright

Some data only exists after a browser has run a page's JavaScript, and some workflows only exist as clicks and forms. Playwright drives real browsers for both, and its async API fits naturally into asyncio programs: every action is an await on a driver process, so one event loop can steer many pages at once. Doing that well comes down to lifecycles and limits, and this section measures them with Playwright 1.63, headless Chromium and Python 3.14 against a local test site of JavaScript-rendered product pages. One browser launched in 90 ms and occupied 174 MiB; a context cost 5 ms and about 20 MiB with a loaded page, so ten pages took 0.70 s with contexts in one browser against 1.36 s with a browser per page. Throughput rose from 2.3 to 12.9 pages per second up to eight concurrent contexts and then stopped, while memory kept climbing to 882 MiB at 32. A 200 ms sleep read a script-rendered price wrongly 9 times in 40; a locator wait was always right and just as fast. Aborting images halved time to the load event, but routing every request through Python cost about 3 ms each. Cancelled jobs that closed their context on a trailing line leaked 6 of 20 contexts; async with leaked none. And a reused page grew Chromium to 906 MiB within 400 pages, where a context per page stayed flat over 1,000.

The parent section, Network I/O & Protocol Handling, covers the HTTP clients and crawlers that should be preferred whenever a page's data is reachable without rendering; this topic is for when it is not. Each guide below measures one part of a browser job's life — launch, concurrency, readiness, routing, cleanup and memory — and gives the pattern that holds up over long runs.

Architectural principles

  • One browser per process, one context per job. Contexts give isolation at a fraction of a browser's cost.
  • Concurrency is a measured plateau, not a maximum. Run the fewest workers that reach peak throughput, within the memory budget.
  • Wait for conditions, never for time. Locators, assertions and responses, not sleeps or networkidle.
  • Route narrowly. Abort what you never need; observe with events; every routed request costs a round trip.
  • Every resource has an owner that releases it on success, failure, timeout and cancellation alike.
How a Playwright service is put together 5 stacked layers. How a Playwright service is put together worker pool fixed size at the throughput plateau; per-host limits browser one per process, recycled every N jobs (145 ms) context per job async with: freed on any exit, ~20 MiB while open routes abort images/fonts; no catch-all continue readiness locator.wait_for / expect / expect_response Each layer corresponds to one guide in this section.

Execution model: an event loop steering a browser

A Playwright program is three processes cooperating: Python, which runs your coroutines; the Playwright driver, a Node.js process that async_playwright() starts; and the browser, which itself is several processes — a browser process plus renderers. Every API call is a message from Python to the driver and on to the browser, awaited until the browser answers. While one page waits for its network or its scripts, the event loop is free to send commands for other pages, which is why a single Python process can drive many pages concurrently, and why blocking work in Python — parsing a large document synchronously — delays every page at once.

Throughput is set by whichever stage saturates first: the target site, the browser's processes, or the CPU running them. On the test site that happened at about eight concurrent contexts, after which more workers only added memory. Memory follows open contexts when pages are closed after each job, and follows pages processed when they are not. Cancellation travels as CancelledError through the awaited calls of a job, so cleanup must be structured — async with — rather than a final line the error skips. The driver cleans up the browser if Python dies: after both SIGTERM and SIGKILL, no Chromium process survived three seconds. Timeouts work the same way at every level: Playwright's per-action default is 30 seconds, a context can set its own default, and an asyncio.timeout around a whole job cancels it at whatever await it has reached, unwinding through the job's async with blocks so its context is closed on the way out.

The measurements behind this section A grid of 6 rows by 3 columns. The measurements behind this section question measured guide browser vs context cost? 90 ms / 174 MiB vs 5 ms / ~20 MiB running Playwright how many concurrent pages? plateau at 8: 12.9 pages/s pooling contexts sleep or wait? sleep 200 ms: 9/40 wrong; locator: 40/40 waiting block or route? 241 -> 108 ms; routing ~3 ms/request routing cleanup on cancellation? trailing close: 6 leaked; async with: 0 cleanup memory over long runs? reused page 906 MiB; context per page flat memory Playwright 1.63, headless Chromium, Python 3.14; shared machine, so rates are approximate.

Pattern catalogue

One browser, a context per job

async with async_playwright() as p:
    browser = await p.chromium.launch()
    async with await browser.new_context() as context:
        page = await context.new_page()
        ...

A context is an isolated profile — its own cookies, storage and cache — created in 5.3 ms, so giving every job its own costs little and removes a whole class of cross-job contamination: a login, a consent cookie or a broken service worker from one job cannot affect the next. Launching a browser per job instead doubled the time for ten pages. See running Playwright with asyncio.

A worker pool sized at the plateau

A fixed number of workers pulling from a queue, each opening a context per URL; calibrate the count by measuring pages per second and PSS at increasing sizes. In the measurements, four workers gave 8.0 pages per second, eight gave 12.9, and sixteen and thirty-two were no faster while peak memory went from 375 to 882 MiB. A fixed pool keeps the number of live contexts exactly equal to the worker count, which is what makes memory predictable. See pooling browser contexts for concurrent pages.

Readiness by condition

await page.goto(url)
await page.locator("#price[data-ready]").wait_for()

Or expect(...) for content and page.expect_response(...) for API data. The selector must match only the final state: on the test page #price existed from the start with the text loading, so the wait targeted the attribute the page set once the API call returned. networkidle was correct but took 901 ms per page against 427 ms for the locator. See waiting for page state without sleeps.

Narrow routes

await context.route("**/*.{png,jpg,gif,webp,woff2}", lambda route: route.abort())

Mocks with route.fulfill for deterministic tests; events, not routes, for observation. Aborting images and CSS cut time to the load event from 241 to 108 ms, while a catch-all route that only continued requests added 72 ms per page — every matched request makes a round trip to a Python handler. Blocking also only helps when the blocked resources sit on the path to the readiness condition. See blocking and mocking requests in Playwright.

Structured cleanup

async with for contexts, asyncio.timeout per job, workers stopped before the browser closes, an init process in containers. The ordering at shutdown matters: closing the browser while workers are still unwinding makes their context closes fail with "has been closed" errors, so the TaskGroup of workers finishes first and the browser closes in the finally after it. See cleaning up Playwright on cancellation.

Bounded memory

PSS-based measurement, pages closed per job, workers sized from a per-page cost, scheduled recycling and a memory gate. Summed RSS reported 321 MiB for a browser whose PSS was 174 MiB, so the metric itself is the first thing to get right; after that, memory that follows open contexts is healthy and memory that follows pages processed is a lifecycle bug. See keeping headless browser memory bounded.

When not to use a browser

Rendering a page is two to three orders of magnitude more expensive than fetching its data: the test pages took about 400 ms each to render and wait for, consumed tens of megabytes while open, and generated 22 requests apiece, where the price itself came from a single JSON call that took 30 ms on the server. Before reaching for Playwright, check whether the data is available more directly: an API the page calls (visible in the browser's network panel), server-rendered HTML, a sitemap or feed, or structured data embedded in the page. Those are better served by an HTTP client and the patterns in Async Web Crawlers, with far higher throughput per core and far lower load on the site.

Browsers earn their cost when the content genuinely requires execution — single-page applications without accessible APIs, pages whose data depends on client-side computation, workflows that need clicks, uploads or form submissions — and when the rendered output itself is the product, as with screenshots and PDFs. A hybrid is common: a browser for login and session establishment, then plain HTTP with the session's cookies for the bulk of the requests, using context.storage_state() to hand the session over.

Is a browser needed for this job? A decision on Where does the data come from with 4 outcomes. Is a browser needed for this job? Where does the data come from? an API or server HTML HTTP client, not a browser ~30 ms vs ~400 ms client-side rendering, interactions Playwright contexts + pool the rendering itself (PDF, screenshot) Playwright output is the product a browser-only login browser, then HTTP with storage_state hybrid Render only what cannot be fetched.

Resource boundaries

  • Browser processes: one per Python process; more only for different engines, proxies, or scaling past one browser's plateau.
  • Open contexts: equal to the worker count; each about 20 MiB on simple pages, far more on heavy sites.
  • Workers: the smaller of the throughput plateau and (memory budget × headroom − browser base) ÷ per-page cost.
  • Per-host concurrency: two or so pages per host, remembering each page is dozens of requests.
  • Job time: an asyncio.timeout per job plus a context default timeout per action.
  • Browser lifetime: recycled every few hundred jobs; replacement launched before the old one is retired.

Integrated production example

A renderer that combines the patterns: one browser per process recycled every N jobs, a fixed worker pool, per-host limits, a context per job with image and font requests aborted at the context, a locator readiness wait, and a timeout per job. Against the local test site it rendered 300 of 300 pages in 24.3 s with eight workers, recycling through three browsers, and left no Chromium processes behind:

import asyncio
import logging
from collections import defaultdict
from urllib.parse import urlsplit

from playwright.async_api import Browser

log = logging.getLogger("render")
BLOCK = "**/*.{png,jpg,jpeg,gif,webp,woff,woff2}"


class Renderer:
    """One browser per process, recycled every N jobs; one context per job."""

    def __init__(self, playwright, workers: int = 8, recycle_every: int = 500,
                 per_host: int = 2, job_timeout: float = 20.0) -> None:
        self.p, self.workers, self.recycle_every = playwright, workers, recycle_every
        self.job_timeout = job_timeout
        self.per_host = defaultdict(lambda: asyncio.Semaphore(per_host))
        self.browser: Browser | None = None
        self.used = 0
        self._lock = asyncio.Lock()
        self._retiring: set[asyncio.Task] = set()

    async def _browser(self) -> Browser:
        async with self._lock:
            if self.browser is None or self.used >= self.recycle_every:
                old, self.browser, self.used = self.browser, await self.p.chromium.launch(), 0
                if old is not None:
                    t = asyncio.create_task(self._retire(old))
                    self._retiring.add(t); t.add_done_callback(self._retiring.discard)
            self.used += 1
            return self.browser

    @staticmethod
    async def _retire(browser: Browser) -> None:
        while browser.contexts:                                  # let in-flight jobs finish
            await asyncio.sleep(0.5)
        await browser.close()

    async def render(self, url: str) -> str | None:
        browser = await self._browser()
        async with self.per_host[urlsplit(url).hostname]:
            try:
                async with asyncio.timeout(self.job_timeout):
                    async with await browser.new_context() as context:
                        await context.route(BLOCK, lambda route: route.abort())
                        page = await context.new_page()
                        await page.goto(url)
                        await page.locator("#price[data-ready]").wait_for()
                        return await page.locator("#price").text_content()
            except TimeoutError:
                log.warning("timed out: %s", url)
            except Exception as exc:                             # one bad page, not the crawl
                log.warning("failed: %s (%s)", url, type(exc).__name__)
            return None

    async def run(self, urls: list[str]) -> dict[str, str | None]:
        queue: asyncio.Queue[str] = asyncio.Queue()
        for u in urls:
            queue.put_nowait(u)
        results: dict[str, str | None] = {}

        async def worker() -> None:
            while not queue.empty():
                url = queue.get_nowait()
                results[url] = await self.render(url)

        try:
            async with asyncio.TaskGroup() as tg:
                for _ in range(self.workers):
                    tg.create_task(worker())
        finally:
            if self.browser is not None:
                await self.browser.close()
            await asyncio.gather(*self._retiring, return_exceptions=True)
        return results

Blocking is set on the context so every page in the job inherits it, and the readiness wait targets the price, which the blocked images never delayed. The retiring browsers are held in a set so their tasks are not garbage-collected mid-close. The test site was on the same machine, so the per-host limit was raised to eight; against real sites, two is a more courteous default.

Diagnostic hook callout

Export, per process: pages per second, jobs timed out, open contexts, browser PSS, and browser recycles. Alert on:

  • Open contexts above the worker count — a job path that does not close its context.
  • Browser PSS trending up with pages processed — reused pages or leaked contexts.
  • Timeouts concentrated on one host — the site changed its markup (the readiness selector no longer matches) or is throttling you.
  • Pages per second falling while workers are busy — the target site or CPU saturated; adding workers will not help.

A calibration run at the start of each large crawl — pages per second and PSS at a few worker counts — catches most of these before they cost a day's run.

Failure modes

Failure mode Root cause Detection Fix
Slow, memory-heavy runs A browser per job Browser processes per job One browser, a context per job
More workers, no more throughput Past the plateau Pages/s vs worker count Stop at the plateau; scale with processes
Wrong or empty values Fixed sleeps, selectors matching loading state Spot checks; loading in results Locators on final state, expect
Slower after adding routes Catch-all routes, or blocking off the critical path Time to readiness with/without Narrow routes; events for observation
Contexts leak on timeouts Close on a trailing line Open contexts above worker count async with
OOM kills after hours Reused pages, no recycling PSS vs pages processed Close per job, recycle, memory gate
Zombie Chromium in containers Python as PID 1 <defunct> processes tini or --init

Frequently Asked Questions

Is Playwright's async API good for concurrent scraping?

Yes: one event loop drove many contexts in one browser, reaching 12.9 pages per second at 8 concurrent contexts in testing. Beyond that point more contexts added memory but no throughput.

How do I avoid memory leaks with Playwright in Python?

Open a context per job with async with so it closes on every exit, never reuse one page for many navigations, and recycle the browser periodically. Reusing pages grew Chromium to 906 MiB within 400 pages in testing; contexts per page stayed flat over 1,000.

Should I use sleep to wait for JavaScript in Playwright?

No. A 200 ms sleep read a script-rendered value wrongly 9 times in 40; a locator wait on the element's final state was always correct and just as fast.

When should I not use a headless browser for scraping?

When the data is available from an API the page calls, from server-rendered HTML, or from a feed: rendering cost about 400 ms per page here, against 30 ms for the underlying JSON call.

Does Playwright leave Chrome processes running if my script crashes?

Not in testing: after SIGTERM and SIGKILL of the Python process, no Chromium process remained three seconds later. In containers, add an init process so exited children are reaped.