Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

How to Handle CAPTCHAs and Bot Detection When Scraping

On this page

How Bot Detection Works

Websites use multiple layers of detection to distinguish human visitors from automated scrapers. Understanding these layers helps you avoid triggering them in the first place.

Layer 1: Request-level signals

The simplest detection methods examine individual HTTP requests:

Signal What it detects How scrapers trigger it
Missing User-Agent Default HTTP libraries send no or library-specific User-Agent Using requests or urllib with default settings
Missing headers Browsers send Accept, Accept-Language, Accept-Encoding and other headers Sending bare requests with only URL
Request frequency Too many requests per second from one IP No rate limiting between requests
Request pattern Sequential URLs, no CSS/JS/image requests Fetching only HTML, skipping assets
TLS fingerprint Library TLS handshake differs from browser Using Python requests (identifiable JA3 fingerprint)

Layer 2: JavaScript challenges

More advanced detection requires JavaScript execution:

Challenge type How it works
Cloudflare Under Attack Mode Serves a JavaScript challenge page; browser must execute JS and set a cookie before accessing the real page
Akamai Bot Manager Injects JavaScript that collects browser environment data and sends it to Akamai's servers for evaluation
DataDome Client-side JavaScript that fingerprints the browser and device
PerimeterX (HUMAN) Behavioural analysis through JavaScript: mouse movements, scroll patterns, typing cadence
Custom JS checks Site-specific scripts that verify browser capabilities (Canvas, WebGL, etc.)

These challenges are invisible to human visitors (the check happens in milliseconds) but block HTTP-only scrapers that cannot execute JavaScript.

Layer 3: Browser fingerprinting

Even with a headless browser, detection systems examine browser properties:

Fingerprint What it reveals
navigator.webdriver Set to true in automated browsers (Selenium, Playwright, Puppeteer)
Chrome plugins Real Chrome has default plugins (PDF viewer, etc.); headless Chrome does not
WebGL renderer Headless environments often show different GPU information
Canvas fingerprint Drawing operations produce different results in headless mode
Screen and window dimensions Headless browsers may report unusual viewport sizes
Timezone and locale Mismatch between IP geolocation and reported timezone
Installed fonts Headless environments have fewer fonts than typical desktops
Permission API Automated browsers respond differently to permission queries
Audio context Audio processing fingerprint differs in headless mode

Layer 4: Behavioural analysis

The most sophisticated systems analyse behaviour over time:

Behaviour Human Bot
Mouse movement Curved, varied, includes micro-movements Absent, straight lines, or mathematically perfect curves
Scroll pattern Variable speed, pauses Absent or uniform
Click timing Variable, includes hover-before-click Instant, no hover
Page dwell time Variable (reading time) Instant or uniformly timed
Navigation pattern Reads content, follows interest-based links Sequential, exhaustive crawling

Types of CAPTCHAs

reCAPTCHA v2 ("I'm not a robot" checkbox)

The checkbox CAPTCHA from Google. If the checkbox alone is insufficient (based on risk scoring), it presents an image challenge (select all images with traffic lights, etc.).

Risk score factors: IP reputation, cookies, browser history with Google services, mouse movement toward the checkbox.

reCAPTCHA v3 (invisible)

No visible challenge. Runs in the background and assigns a score (0.0 to 1.0) based on user behaviour. The website decides what score threshold triggers a block or additional verification.

How it works: A JavaScript snippet runs on the page, observes user behaviour and sends data to Google. The site receives a score and decides how to act.

hCaptcha

An alternative to reCAPTCHA that presents image-based challenges. Widely adopted after Google began charging for reCAPTCHA in some cases.

Cloudflare Turnstile

Cloudflare's managed challenge that replaces traditional CAPTCHAs. Uses a combination of browser challenges, proof-of-work and machine learning to verify visitors without visible challenges in most cases.

FunCaptcha (Arkose Labs)

Interactive challenges (rotate a 3D object, match an image) designed to be harder for automated solvers.

Prevention: Avoiding Detection

The most effective approach to bot detection is not solving CAPTCHAs, but avoiding triggering them in the first place.

Request-level best practices

Set complete headers:

headers = {
    "User-Agent": (
        "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
        "AppleWebKit/537.36 (KHTML, like Gecko) "
        "Chrome/120.0.0.0 Safari/537.36"
    ),
    "Accept": (
        "text/html,application/xhtml+xml,application/xml;"
        "q=0.9,image/webp,*/*;q=0.8"
    ),
    "Accept-Language": "en-US,en;q=0.9",
    "Accept-Encoding": "gzip, deflate, br",
    "Connection": "keep-alive",
    "Upgrade-Insecure-Requests": "1",
}

Maintain sessions: Use requests.Session() to persist cookies and connection state across requests.

Respect rate limits: Add random delays between requests (2-5 seconds minimum). Never hit the same domain more than once per second.

Browser-level best practices (headless browsers)

Patch navigator.webdriver:

// Playwright
await page.addInitScript(() => {
    Object.defineProperty(navigator, "webdriver", {
        get: () => undefined,
    });
});

Set realistic viewport:

# Playwright
context = browser.new_context(
    viewport={"width": 1920, "height": 1080},
    screen={"width": 1920, "height": 1080},
)

Use stealth plugins:

  • Puppeteer: puppeteer-extra-plugin-stealth patches multiple detection vectors.
  • Selenium: undetected-chromedriver patches ChromeDriver to avoid detection.
  • Playwright: No official stealth plugin, but the community maintains patches for common detection vectors.

Behavioural best practices

  • Do not scrape every page on a site. Target only the pages that matter (contact pages, team pages).
  • Mix up your request pattern. Do not visit pages in alphabetical or sequential order.
  • Visit the homepage first before navigating to internal pages.
  • Avoid scraping at perfectly regular intervals.

Handling CAPTCHAs When They Appear

Option 1: Slow down and retry

Many CAPTCHAs are triggered by rate limits. When you encounter one:

  1. Stop making requests to that domain.
  2. Wait 5-10 minutes.
  3. Clear cookies and use a different IP.
  4. Retry with longer delays between requests.

Option 2: Use a different approach

If a website consistently blocks scraping:

  • Check for an API. Many websites with aggressive bot protection offer a public API for structured data access.
  • Check for a data feed. Some sites publish RSS, sitemap or data export options.
  • Use Email Extractor. If you can access the page in your browser, copy the page content and paste it into Email Extractor to extract email addresses without scraping.
  • Use the Email Extractor browser extension. The browser extension extracts emails from pages you are already viewing in Chrome or Edge, without automated scraping.

Option 3: CAPTCHA solving services

Third-party services solve CAPTCHAs using human workers or machine learning:

Service type How it works Typical speed
Human solving Sends CAPTCHA image to a human worker who solves it 10-30 seconds
AI solving Machine learning models trained on specific CAPTCHA types 1-10 seconds
Browser-based Manages a browser session that handles challenges automatically Variable

Costs: Human solving typically costs $1-3 per 1,000 CAPTCHAs. AI solving costs less but has lower accuracy on complex challenges.

Ethical and legal considerations: Using CAPTCHA solving services circumvents a website's intentional access control. This may violate the website's terms of service and could raise legal issues under computer fraud and abuse laws in some jurisdictions. Evaluate the legal implications for your specific situation before using these services.

Bot Detection by Provider

Cloudflare

Detection level: Moderate to high (depends on the website's configuration).

How it works: Cloudflare sits between the visitor and the website. It evaluates each request based on IP reputation, browser fingerprint, challenge completion and behavioural signals.

Challenge types:

  • JavaScript challenge (5-second page).
  • Managed challenge (Turnstile).
  • Interactive challenge (CAPTCHA when risk is high).

Common triggers:

  • Datacenter IP addresses.
  • Missing or unusual User-Agent.
  • High request rate.
  • Failed JavaScript challenge.

Akamai

Detection level: High.

How it works: Akamai's Bot Manager uses sensor data collected by client-side JavaScript, device fingerprinting and behavioural analysis.

Key detection: Akamai's sensor script generates a token that must be included in subsequent requests. Without executing the script and passing the token, requests are blocked.

PerimeterX (HUMAN)

Detection level: Very high.

How it works: Collects granular behavioural data (mouse movements, keystrokes, scroll patterns) and uses machine learning to distinguish human from automated behaviour. Specifically designed to defeat headless browsers.

When Not to Scrape

Sometimes the correct response to bot detection is to not scrape:

  • The data is available through an API. Use the API instead.
  • The data is available for purchase. Data providers may have the data you need at a reasonable cost.
  • The website explicitly prohibits scraping. Respect terms of service, especially for sites where you have an account.
  • The data is not public. Scraping data behind a login may violate computer access laws.
  • The effort exceeds the value. If a site's bot protection requires extensive engineering to bypass, the cost may exceed the value of the data.

For extracting emails from files and documents you already have access to, use Email Extractor instead of scraping. Upload PDFs, spreadsheets, text files and other supported formats directly, and the tool extracts and deduplicates email addresses in your browser.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)