Headless Browsers for Scraping: Puppeteer vs Playwright vs Selenium
On this page
When You Need a Headless Browser
Standard HTTP scraping (requests + HTML parsing) works when the data you need is in the initial HTML response. A headless browser is necessary when:
- The page renders content with JavaScript (single-page applications, React/Vue/Angular sites).
- Data loads asynchronously after the initial page load (infinite scroll, lazy loading, API-driven content).
- You need to interact with the page (click buttons, fill forms, navigate pagination).
- The site checks for browser-like behaviour (cookies, JavaScript execution, user agent features).
If the data is available in the raw HTML or through a public API, a headless browser adds unnecessary overhead. Check the page source and network tab first.
The Three Frameworks
Puppeteer
Maintained by: Google Chrome team. Language: JavaScript/TypeScript (Node.js). Default browser: Chromium (Chrome). Protocol: Chrome DevTools Protocol (CDP).
Puppeteer was built specifically for Chrome automation. It controls Chromium through the DevTools Protocol, the same protocol Chrome DevTools uses internally. This gives it deep, native access to Chrome features.
Architecture:
Your script → Puppeteer API → CDP (WebSocket) → Chromium
Playwright
Maintained by: Microsoft. Language: JavaScript/TypeScript, Python, Java, C#. Default browser: Chromium, Firefox, WebKit (all three). Protocol: CDP for Chromium, custom protocols for Firefox and WebKit.
Playwright was created by former Puppeteer team members at Microsoft. It extends the Puppeteer model to support multiple browsers and adds features like auto-waiting, network interception and test isolation.
Architecture:
Your script → Playwright API → Browser-specific protocol → Chromium/Firefox/WebKit
Selenium
Maintained by: Open-source community (Selenium project). Language: Java, Python, JavaScript, C#, Ruby, Kotlin. Default browser: None (install any supported browser separately). Protocol: WebDriver (W3C standard).
Selenium is the oldest browser automation framework. It uses the WebDriver protocol, which is a W3C standard implemented by every major browser vendor. This standardisation is both its strength and limitation.
Architecture:
Your script → Selenium client → WebDriver HTTP API → Browser driver → Browser
Feature Comparison
| Feature | Puppeteer | Playwright | Selenium |
|---|---|---|---|
| Language support | JavaScript/TypeScript | JS/TS, Python, Java, C# | Java, Python, JS, C#, Ruby, Kotlin |
| Browser support | Chromium (experimental Firefox) | Chromium, Firefox, WebKit | Chrome, Firefox, Safari, Edge, IE |
| Installation | npm install puppeteer (downloads Chromium) | npm install playwright (downloads all 3 browsers) | pip install selenium + separate browser driver |
| Auto-waiting | Manual (waitForSelector, waitForNavigation) | Built-in (auto-waits for elements to be actionable) | Manual (explicit/implicit waits) |
| Network interception | Yes (CDP) | Yes (built-in API) | Limited (requires proxy) |
| Multiple browser contexts | Yes | Yes (isolated by default) | Requires separate driver instances |
| Mobile emulation | Yes | Yes | Limited |
| PDF generation | Yes | Yes | No (requires workarounds) |
| Screenshot | Yes | Yes (full page, element, clip) | Yes |
| File download handling | Manual configuration | Built-in API | Manual configuration |
| iframe support | switchTo frame | Built-in locator support | switchTo frame |
| Shadow DOM | Yes (CDP piercing) | Built-in locator support | Limited |
| Protocol | CDP (Chrome-native) | CDP + custom (multi-browser) | WebDriver (W3C standard) |
| Parallel execution | Manual (cluster libraries) | Built-in (browser contexts) | Selenium Grid |
Performance for Scraping
Startup time
Playwright and Puppeteer both launch a browser process and connect over a local WebSocket. Selenium launches a browser driver process, then the browser, and communicates over HTTP. This extra layer adds latency.
| Framework | Typical browser launch time | Per-page overhead |
|---|---|---|
| Puppeteer | 500-1500 ms | Low (direct CDP) |
| Playwright | 500-1500 ms | Low (direct protocol) |
| Selenium | 1000-3000 ms | Higher (HTTP round-trips to WebDriver) |
Memory usage
Each browser instance consumes significant memory. For scraping at scale, memory is often the limiting factor.
| Approach | Memory per instance |
|---|---|
| Chromium (headless) | 80-200 MB per tab (varies by page complexity) |
| Firefox (headless) | 100-250 MB per tab |
| WebKit (headless) | 60-150 MB per tab |
Playwright's WebKit option uses less memory than Chromium, which can matter at scale.
Parallelisation
| Framework | Parallel approach | Notes |
|---|---|---|
| Puppeteer | puppeteer-cluster library, or manual process management | No built-in parallelism |
| Playwright | Browser contexts (lightweight, isolated) | Contexts share a single browser process, reducing memory |
| Selenium | Selenium Grid (separate machines or containers) | Heavy infrastructure for parallelism |
Playwright's browser contexts are the most efficient approach. Each context has its own cookies, storage and cache, but shares the browser process. This allows dozens of parallel scraping sessions in one browser instance.
Scraping-Specific Features
Waiting for content
Dynamic pages require waiting for JavaScript to render content before extracting data.
Puppeteer:
await page.goto("https://example.com/directory");
await page.waitForSelector(".contact-card");
const emails = await page.$$eval(".contact-card .email",
els => els.map(el => el.textContent)
);
Playwright:
page.goto("https://example.com/directory")
# Playwright auto-waits for elements to be visible and stable
emails = page.locator(".contact-card .email").all_text_contents()
Selenium:
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
driver.get("https://example.com/directory")
WebDriverWait(driver, 10).until(
EC.presence_of_all_elements_located(
(By.CSS_SELECTOR, ".contact-card")
)
)
emails = [el.text for el in driver.find_elements(
By.CSS_SELECTOR, ".contact-card .email"
)]
Playwright's auto-waiting reduces the most common source of flaky scraping scripts: timing issues where the script tries to read content before JavaScript has rendered it.
Network interception
Intercepting network requests lets you block unnecessary resources (images, stylesheets, analytics) to speed up page loads, or capture API responses directly.
Puppeteer:
await page.setRequestInterception(true);
page.on("request", request => {
if (["image", "stylesheet", "font"].includes(request.resourceType())) {
request.abort();
} else {
request.continue();
}
});
Playwright:
def handle_route(route):
if route.request.resource_type in ["image", "stylesheet", "font"]:
route.abort()
else:
route.fallback()
page.route("**/*", handle_route)
Selenium: Native network interception is limited. Most Selenium-based scrapers use a proxy (like BrowserMob Proxy or mitmproxy) to intercept and modify traffic.
Capturing API responses
Many modern websites load data from internal APIs. Intercepting these responses can be more efficient than parsing the rendered DOM.
Playwright:
def handle_response(response):
if "/api/contacts" in response.url:
data = response.json()
# Process the structured data directly
page.on("response", handle_response)
page.goto("https://example.com/directory")
This approach often yields cleaner data than DOM parsing, because the API response contains structured JSON rather than HTML that needs to be scraped.
Handling pagination
Puppeteer:
let hasNextPage = true;
while (hasNextPage) {
// Extract data from current page
const data = await page.evaluate(() => {
return Array.from(document.querySelectorAll(".email"))
.map(el => el.textContent);
});
allEmails.push(...data);
// Check for and click next button
const nextButton = await page.$(".pagination .next:not(.disabled)");
if (nextButton) {
await nextButton.click();
await page.waitForNavigation();
} else {
hasNextPage = false;
}
}
Playwright:
while True:
emails = page.locator(".email").all_text_contents()
all_emails.extend(emails)
next_button = page.locator(".pagination .next:not(.disabled)")
if next_button.count() > 0:
next_button.click()
page.wait_for_load_state("networkidle")
else:
break
Anti-Detection Considerations
Websites use various methods to detect and block automated browsers. Each framework has different characteristics in this regard.
Browser fingerprinting
| Detection method | Puppeteer | Playwright | Selenium |
|---|---|---|---|
| navigator.webdriver property | Set to true by default | Set to true by default | Set to true by default |
| Headless detection (missing Chrome plugins) | Detectable | Detectable | Detectable |
| WebDriver protocol artifacts | None (uses CDP) | None (uses CDP) | Detectable (WebDriver endpoints) |
| Automation extension | Present by default | Not present | Present by default |
| User agent string | Default includes "HeadlessChrome" | Default includes "HeadlessChrome" | Normal browser user agent |
Stealth approaches
Puppeteer: The puppeteer-extra-plugin-stealth package patches many detection vectors automatically.
const puppeteer = require("puppeteer-extra");
const StealthPlugin = require("puppeteer-extra-plugin-stealth");
puppeteer.use(StealthPlugin());
Playwright: No official stealth plugin, but you can patch individual detection vectors through browser context options and JavaScript injection.
Selenium: undetected-chromedriver (Python) patches ChromeDriver to avoid common detection methods.
Best practices for avoiding blocks
- Set a realistic user agent. Replace the default headless user agent with a current desktop browser string.
- Add realistic viewport size. Default headless viewports (800x600) are a detection signal.
- Respect rate limits. Add delays between requests. Random delays (2-5 seconds) look more natural than fixed intervals.
- Rotate IP addresses. Residential or datacenter proxy rotation avoids IP-based rate limiting.
- Handle cookies. Accept cookie banners and maintain session cookies across requests.
Choosing a Framework
Choose Puppeteer when
- You only need Chrome/Chromium.
- You work in JavaScript/TypeScript.
- You want the most direct Chrome API access (CDP).
- You need Chrome-specific features (PDF generation, performance tracing).
- You have existing Puppeteer code to maintain.
Choose Playwright when
- You need multi-browser support (Chrome, Firefox, WebKit).
- You want built-in auto-waiting (reduces flaky scripts).
- You need efficient parallel scraping (browser contexts).
- You work in Python, Java or C# (not just JavaScript).
- You want built-in network interception without a proxy.
- You are starting a new scraping project (most modern API design).
Choose Selenium when
- You need to support older or niche browsers.
- You have existing Selenium infrastructure (Grid, test suites).
- You need the W3C WebDriver standard for compliance.
- Your team knows Selenium and switching cost is high.
- You need Ruby or Kotlin support.
For new scraping projects, Playwright is generally the strongest choice. Its auto-waiting, browser context isolation and multi-language support address the most common scraping pain points. Puppeteer remains strong for Chrome-only use cases where you want the thinnest abstraction over CDP. Selenium's strength is ecosystem maturity and language breadth, but its HTTP-based architecture adds overhead that matters at scraping scale.
Processing Scraped Data
Once you have extracted text or downloaded files from scraped pages, upload the results to Email Extractor to isolate email addresses:
- Save scraped page content or downloaded files locally.
- Go to Email Extractor.
- Select "Text and files."
- Upload the saved files (supported formats include TXT, HTML, CSV, JSON, PDF, DOCX and others, up to 25 MB per file).
- Click "Extract emails."
The tool identifies all email addresses in the uploaded content, removes duplicates using case-insensitive matching and lets you download the clean list as TXT, CSV or CSV with sources.