Running Playwright with asyncio¶
Playwright drives real browsers — Chromium, Firefox, WebKit — through a driver process, and its async_api exposes every operation as a coroutine, so a single asyncio program can drive many pages concurrently. The structure that works is one browser per process, one context per independent session, and pages inside contexts; the costs that justify it were measured on Python 3.14 with Playwright 1.63 and its headless Chromium shell, against a local test site whose product pages render a price with JavaScript after a 50–400 ms delay. Launching the browser took 90 ms and 174 MiB (proportional set size across Chromium's processes). A new context took 5.3 ms, a new page 42 ms, and each context with a loaded page added about 20 MiB. Loading ten pages concurrently with a browser launched per page took 1.36 s; with one browser and a context per page, 0.70 s. This guide sets up that structure correctly.
Prerequisites¶
- Python 3.9+,
pip install playwrightandplaywright install chromium(tested with Playwright 1.63). - Task groups and cancellation, from structured concurrency with asyncio.TaskGroup.
- The topic overview, Browser Automation.
1. Start Playwright and one browser per process¶
async_playwright() starts the driver; the browser is launched once and shared by everything in the process:
import asyncio
from playwright.async_api import async_playwright
async def main() -> None:
async with async_playwright() as p:
browser = await p.chromium.launch() # headless by default
try:
await run_jobs(browser)
finally:
await browser.close()
asyncio.run(main())
Measured: the launch took 90 ms and the idle browser occupied 174 MiB of PSS across its processes. The async with block starts and stops the Playwright driver — a Node.js process that talks to the browser — and closing it also stops any browser it launched. Use PSS rather than summed RSS when measuring Chromium: it runs several processes that share memory, and summed RSS counted the same browser at 321 MiB.
Verify: ps shows one browser process tree per Python process, not one per job.
2. Use a context per independent job¶
A BrowserContext is an isolated profile: its own cookies, local storage, cache and permissions. Contexts are cheap, so give each independent job its own and close it when the job ends:
async def price_of(browser, product_id: int) -> str:
async with await browser.new_context() as context: # isolated session
page = await context.new_page()
await page.goto(f"https://shop.example/p/{product_id}")
await page.locator("#price[data-ready]").wait_for()
return await page.locator("#price").text_content()
Measured: new_context() took 5.3 ms and new_page() 42 ms; with twenty contexts open, each holding a loaded page, PSS rose from 174 MiB to 623 MiB — about 20 MiB per context — and fell back to 209 MiB when they closed. Isolation is the main reason to prefer contexts over reusing one page: a login, a cookie or a broken service worker from one job cannot leak into the next. async with closes the context even when the job fails or is cancelled, which the cleanup guide shows matters.
Verify: after a batch of jobs, len(browser.contexts) is back to zero.
3. Run pages concurrently, not browsers¶
Concurrency in Playwright means several contexts in one browser, driven by concurrent tasks:
async def main() -> None:
async with async_playwright() as p:
browser = await p.chromium.launch()
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(price_of(browser, i)) for i in range(10)]
print([t.result() for t in tasks])
await browser.close()
Measured for ten product pages: 0.70 s with one browser and a context per page, against 1.36 s when each page launched its own browser — the extra launches, and ten copies of the browser's base memory, bought nothing. A browser per job makes sense only when jobs need different browser-level settings (a proxy per job, a different browser engine) or must be isolated from each other's crashes. How many contexts to run at once is the subject of pooling browser contexts for concurrent pages.
Verify: doubling the number of concurrent jobs does not double the number of browser processes.
4. Keep blocking work out of page tasks¶
Every Playwright call is an await on the driver, so the event loop is free while the browser works. The risk is the code around it — parsing large HTML with a synchronous parser, writing screenshots, calling a blocking API with the results:
from selectolax.parser import HTMLParser
async def scrape(browser, url: str) -> list[str]:
async with await browser.new_context() as context:
page = await context.new_page()
await page.goto(url)
html = await page.content()
# parsing a large document is CPU work: keep it off the event loop
return await asyncio.to_thread(lambda: [n.text() for n in HTMLParser(html).css("h2")])
While the parser runs on the loop, every other page's events — navigation responses, console messages, route handlers — wait, and Playwright's own timeouts keep counting. Prefer extracting data in the browser with locators (await page.locator("h2").all_text_contents()), which moves the work into the browser process, and move unavoidable CPU work to a thread or process as in offloading CPU work with loop.run_in_executor.
Verify: event-loop lag stays low while many pages are being processed.
5. Choose headless mode and channel deliberately¶
chromium.launch() uses Playwright's bundled headless Chromium by default. A few launch options change behaviour in ways that matter for automation at scale:
browser = await p.chromium.launch(
headless=True, # default; no display needed
args=["--disable-dev-shm-usage"], # in containers with a small /dev/shm
timeout=30_000, # launch timeout, ms
)
context = await browser.new_context(
viewport={"width": 1280, "height": 800},
user_agent="example-bot/1.0 (+https://example.com/bot)",
locale="en-US",
)
context.set_default_timeout(15_000) # every action and wait in this context
Containers often mount a small /dev/shm, which Chromium uses for shared memory between its processes; --disable-dev-shm-usage makes it use /tmp instead and avoids crashes under load. Setting a default timeout per context bounds every wait and action — Playwright's default is 30 seconds — so a stuck page fails its job instead of holding a task. An honest user agent and the crawling etiquette in respecting robots.txt and crawl-delay asynchronously apply to browser-driven crawls as much as to HTTP clients.
Verify: a page that never finishes loading fails within the context's default timeout.
Verification¶
A Playwright program is structured well when:
- One browser serves the whole process, launched once inside
async_playwright(). - Each independent job has its own context, closed with
async with. - Concurrency comes from tasks and contexts, not from launching browsers.
- CPU-heavy processing of page content is off the event loop, and contexts have default timeouts.
Diagnostic Hook: chart open contexts and Chromium PSS over time. Open contexts should rise and fall with the number of jobs in flight, and memory should follow; a context count that only rises means some code path never closes them.
Pitfalls & edge cases¶
- A browser per job. Measured: twice as slow as contexts for ten pages, with a browser's memory each.
- Summed RSS for Chromium. It double-counts shared memory; use PSS.
- Synchronous parsing in page tasks. It stalls every other page on the loop.
- Default 30 s timeouts. Set a context default that matches your job budget.
Frequently Asked Questions¶
How do I use Playwright with asyncio?
Use playwright.async_api: async with async_playwright() as p, launch one browser with await p.chromium.launch(), and run each job in its own context with async with await browser.new_context() as ctx. Every operation is awaited.
Should I launch a browser per task or share one?
Share one and use a context per task: ten pages took 0.70 s with one browser and contexts, against 1.36 s with a browser per page, and a context added about 20 MiB where a browser costs about 174 MiB.
What is a BrowserContext in Playwright?
An isolated browser profile within one browser — its own cookies, storage and cache. Creating one took about 5 ms in testing, so a context per job is affordable.
How much memory does headless Chromium use?
About 174 MiB of PSS for the idle browser and about 20 MiB more per context with a loaded page in testing; measure PSS rather than summed RSS, which double-counts shared memory.
Related¶
- Browser Automation — up to the topic overview.
- Pooling browser contexts for concurrent pages — how many at once.
- Network I/O & Protocol Handling — the section overview.