Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Headless Browsers for Scraping: Puppeteer vs Playwright vs Selenium

On this page

When You Need a Headless Browser

Standard HTTP scraping (requests + HTML parsing) works when the data you need is in the initial HTML response. A headless browser is necessary when:

  • The page renders content with JavaScript (single-page applications, React/Vue/Angular sites).
  • Data loads asynchronously after the initial page load (infinite scroll, lazy loading, API-driven content).
  • You need to interact with the page (click buttons, fill forms, navigate pagination).
  • The site checks for browser-like behaviour (cookies, JavaScript execution, user agent features).

If the data is available in the raw HTML or through a public API, a headless browser adds unnecessary overhead. Check the page source and network tab first.

The Three Frameworks

Puppeteer

Maintained by: Google Chrome team. Language: JavaScript/TypeScript (Node.js). Default browser: Chromium (Chrome). Protocol: Chrome DevTools Protocol (CDP).

Puppeteer was built specifically for Chrome automation. It controls Chromium through the DevTools Protocol, the same protocol Chrome DevTools uses internally. This gives it deep, native access to Chrome features.

Architecture:

Your script → Puppeteer API → CDP (WebSocket) → Chromium

Playwright

Maintained by: Microsoft. Language: JavaScript/TypeScript, Python, Java, C#. Default browser: Chromium, Firefox, WebKit (all three). Protocol: CDP for Chromium, custom protocols for Firefox and WebKit.

Playwright was created by former Puppeteer team members at Microsoft. It extends the Puppeteer model to support multiple browsers and adds features like auto-waiting, network interception and test isolation.

Architecture:

Your script → Playwright API → Browser-specific protocol → Chromium/Firefox/WebKit

Selenium

Maintained by: Open-source community (Selenium project). Language: Java, Python, JavaScript, C#, Ruby, Kotlin. Default browser: None (install any supported browser separately). Protocol: WebDriver (W3C standard).

Selenium is the oldest browser automation framework. It uses the WebDriver protocol, which is a W3C standard implemented by every major browser vendor. This standardisation is both its strength and limitation.

Architecture:

Your script → Selenium client → WebDriver HTTP API → Browser driver → Browser

Feature Comparison

Feature Puppeteer Playwright Selenium
Language support JavaScript/TypeScript JS/TS, Python, Java, C# Java, Python, JS, C#, Ruby, Kotlin
Browser support Chromium (experimental Firefox) Chromium, Firefox, WebKit Chrome, Firefox, Safari, Edge, IE
Installation npm install puppeteer (downloads Chromium) npm install playwright (downloads all 3 browsers) pip install selenium + separate browser driver
Auto-waiting Manual (waitForSelector, waitForNavigation) Built-in (auto-waits for elements to be actionable) Manual (explicit/implicit waits)
Network interception Yes (CDP) Yes (built-in API) Limited (requires proxy)
Multiple browser contexts Yes Yes (isolated by default) Requires separate driver instances
Mobile emulation Yes Yes Limited
PDF generation Yes Yes No (requires workarounds)
Screenshot Yes Yes (full page, element, clip) Yes
File download handling Manual configuration Built-in API Manual configuration
iframe support switchTo frame Built-in locator support switchTo frame
Shadow DOM Yes (CDP piercing) Built-in locator support Limited
Protocol CDP (Chrome-native) CDP + custom (multi-browser) WebDriver (W3C standard)
Parallel execution Manual (cluster libraries) Built-in (browser contexts) Selenium Grid

Performance for Scraping

Startup time

Playwright and Puppeteer both launch a browser process and connect over a local WebSocket. Selenium launches a browser driver process, then the browser, and communicates over HTTP. This extra layer adds latency.

Framework Typical browser launch time Per-page overhead
Puppeteer 500-1500 ms Low (direct CDP)
Playwright 500-1500 ms Low (direct protocol)
Selenium 1000-3000 ms Higher (HTTP round-trips to WebDriver)

Memory usage

Each browser instance consumes significant memory. For scraping at scale, memory is often the limiting factor.

Approach Memory per instance
Chromium (headless) 80-200 MB per tab (varies by page complexity)
Firefox (headless) 100-250 MB per tab
WebKit (headless) 60-150 MB per tab

Playwright's WebKit option uses less memory than Chromium, which can matter at scale.

Parallelisation

Framework Parallel approach Notes
Puppeteer puppeteer-cluster library, or manual process management No built-in parallelism
Playwright Browser contexts (lightweight, isolated) Contexts share a single browser process, reducing memory
Selenium Selenium Grid (separate machines or containers) Heavy infrastructure for parallelism

Playwright's browser contexts are the most efficient approach. Each context has its own cookies, storage and cache, but shares the browser process. This allows dozens of parallel scraping sessions in one browser instance.

Scraping-Specific Features

Waiting for content

Dynamic pages require waiting for JavaScript to render content before extracting data.

Puppeteer:

await page.goto("https://example.com/directory");
await page.waitForSelector(".contact-card");
const emails = await page.$$eval(".contact-card .email", 
  els => els.map(el => el.textContent)
);

Playwright:

page.goto("https://example.com/directory")
# Playwright auto-waits for elements to be visible and stable
emails = page.locator(".contact-card .email").all_text_contents()

Selenium:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver.get("https://example.com/directory")
WebDriverWait(driver, 10).until(
    EC.presence_of_all_elements_located(
        (By.CSS_SELECTOR, ".contact-card")
    )
)
emails = [el.text for el in driver.find_elements(
    By.CSS_SELECTOR, ".contact-card .email"
)]

Playwright's auto-waiting reduces the most common source of flaky scraping scripts: timing issues where the script tries to read content before JavaScript has rendered it.

Network interception

Intercepting network requests lets you block unnecessary resources (images, stylesheets, analytics) to speed up page loads, or capture API responses directly.

Puppeteer:

await page.setRequestInterception(true);
page.on("request", request => {
  if (["image", "stylesheet", "font"].includes(request.resourceType())) {
    request.abort();
  } else {
    request.continue();
  }
});

Playwright:

def handle_route(route):
    if route.request.resource_type in ["image", "stylesheet", "font"]:
        route.abort()
    else:
        route.fallback()

page.route("**/*", handle_route)

Selenium: Native network interception is limited. Most Selenium-based scrapers use a proxy (like BrowserMob Proxy or mitmproxy) to intercept and modify traffic.

Capturing API responses

Many modern websites load data from internal APIs. Intercepting these responses can be more efficient than parsing the rendered DOM.

Playwright:

def handle_response(response):
    if "/api/contacts" in response.url:
        data = response.json()
        # Process the structured data directly

page.on("response", handle_response)
page.goto("https://example.com/directory")

This approach often yields cleaner data than DOM parsing, because the API response contains structured JSON rather than HTML that needs to be scraped.

Handling pagination

Puppeteer:

let hasNextPage = true;
while (hasNextPage) {
  // Extract data from current page
  const data = await page.evaluate(() => {
    return Array.from(document.querySelectorAll(".email"))
      .map(el => el.textContent);
  });
  allEmails.push(...data);
  
  // Check for and click next button
  const nextButton = await page.$(".pagination .next:not(.disabled)");
  if (nextButton) {
    await nextButton.click();
    await page.waitForNavigation();
  } else {
    hasNextPage = false;
  }
}

Playwright:

while True:
    emails = page.locator(".email").all_text_contents()
    all_emails.extend(emails)
    
    next_button = page.locator(".pagination .next:not(.disabled)")
    if next_button.count() > 0:
        next_button.click()
        page.wait_for_load_state("networkidle")
    else:
        break

Anti-Detection Considerations

Websites use various methods to detect and block automated browsers. Each framework has different characteristics in this regard.

Browser fingerprinting

Detection method Puppeteer Playwright Selenium
navigator.webdriver property Set to true by default Set to true by default Set to true by default
Headless detection (missing Chrome plugins) Detectable Detectable Detectable
WebDriver protocol artifacts None (uses CDP) None (uses CDP) Detectable (WebDriver endpoints)
Automation extension Present by default Not present Present by default
User agent string Default includes "HeadlessChrome" Default includes "HeadlessChrome" Normal browser user agent

Stealth approaches

Puppeteer: The puppeteer-extra-plugin-stealth package patches many detection vectors automatically.

const puppeteer = require("puppeteer-extra");
const StealthPlugin = require("puppeteer-extra-plugin-stealth");
puppeteer.use(StealthPlugin());

Playwright: No official stealth plugin, but you can patch individual detection vectors through browser context options and JavaScript injection.

Selenium: undetected-chromedriver (Python) patches ChromeDriver to avoid common detection methods.

Best practices for avoiding blocks

  1. Set a realistic user agent. Replace the default headless user agent with a current desktop browser string.
  2. Add realistic viewport size. Default headless viewports (800x600) are a detection signal.
  3. Respect rate limits. Add delays between requests. Random delays (2-5 seconds) look more natural than fixed intervals.
  4. Rotate IP addresses. Residential or datacenter proxy rotation avoids IP-based rate limiting.
  5. Handle cookies. Accept cookie banners and maintain session cookies across requests.

Choosing a Framework

Choose Puppeteer when

  • You only need Chrome/Chromium.
  • You work in JavaScript/TypeScript.
  • You want the most direct Chrome API access (CDP).
  • You need Chrome-specific features (PDF generation, performance tracing).
  • You have existing Puppeteer code to maintain.

Choose Playwright when

  • You need multi-browser support (Chrome, Firefox, WebKit).
  • You want built-in auto-waiting (reduces flaky scripts).
  • You need efficient parallel scraping (browser contexts).
  • You work in Python, Java or C# (not just JavaScript).
  • You want built-in network interception without a proxy.
  • You are starting a new scraping project (most modern API design).

Choose Selenium when

  • You need to support older or niche browsers.
  • You have existing Selenium infrastructure (Grid, test suites).
  • You need the W3C WebDriver standard for compliance.
  • Your team knows Selenium and switching cost is high.
  • You need Ruby or Kotlin support.

For new scraping projects, Playwright is generally the strongest choice. Its auto-waiting, browser context isolation and multi-language support address the most common scraping pain points. Puppeteer remains strong for Chrome-only use cases where you want the thinnest abstraction over CDP. Selenium's strength is ecosystem maturity and language breadth, but its HTTP-based architecture adds overhead that matters at scraping scale.

Processing Scraped Data

Once you have extracted text or downloaded files from scraped pages, upload the results to Email Extractor to isolate email addresses:

  1. Save scraped page content or downloaded files locally.
  2. Go to Email Extractor.
  3. Select "Text and files."
  4. Upload the saved files (supported formats include TXT, HTML, CSV, JSON, PDF, DOCX and others, up to 25 MB per file).
  5. Click "Extract emails."

The tool identifies all email addresses in the uploaded content, removes duplicates using case-insensitive matching and lets you download the clean list as TXT, CSV or CSV with sources.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)