How to Handle CAPTCHAs and Bot Detection When Scraping
On this page
How Bot Detection Works
Websites use multiple layers of detection to distinguish human visitors from automated scrapers. Understanding these layers helps you avoid triggering them in the first place.
Layer 1: Request-level signals
The simplest detection methods examine individual HTTP requests:
| Signal | What it detects | How scrapers trigger it |
|---|---|---|
| Missing User-Agent | Default HTTP libraries send no or library-specific User-Agent | Using requests or urllib with default settings |
| Missing headers | Browsers send Accept, Accept-Language, Accept-Encoding and other headers | Sending bare requests with only URL |
| Request frequency | Too many requests per second from one IP | No rate limiting between requests |
| Request pattern | Sequential URLs, no CSS/JS/image requests | Fetching only HTML, skipping assets |
| TLS fingerprint | Library TLS handshake differs from browser | Using Python requests (identifiable JA3 fingerprint) |
Layer 2: JavaScript challenges
More advanced detection requires JavaScript execution:
| Challenge type | How it works |
|---|---|
| Cloudflare Under Attack Mode | Serves a JavaScript challenge page; browser must execute JS and set a cookie before accessing the real page |
| Akamai Bot Manager | Injects JavaScript that collects browser environment data and sends it to Akamai's servers for evaluation |
| DataDome | Client-side JavaScript that fingerprints the browser and device |
| PerimeterX (HUMAN) | Behavioural analysis through JavaScript: mouse movements, scroll patterns, typing cadence |
| Custom JS checks | Site-specific scripts that verify browser capabilities (Canvas, WebGL, etc.) |
These challenges are invisible to human visitors (the check happens in milliseconds) but block HTTP-only scrapers that cannot execute JavaScript.
Layer 3: Browser fingerprinting
Even with a headless browser, detection systems examine browser properties:
| Fingerprint | What it reveals |
|---|---|
| navigator.webdriver | Set to true in automated browsers (Selenium, Playwright, Puppeteer) |
| Chrome plugins | Real Chrome has default plugins (PDF viewer, etc.); headless Chrome does not |
| WebGL renderer | Headless environments often show different GPU information |
| Canvas fingerprint | Drawing operations produce different results in headless mode |
| Screen and window dimensions | Headless browsers may report unusual viewport sizes |
| Timezone and locale | Mismatch between IP geolocation and reported timezone |
| Installed fonts | Headless environments have fewer fonts than typical desktops |
| Permission API | Automated browsers respond differently to permission queries |
| Audio context | Audio processing fingerprint differs in headless mode |
Layer 4: Behavioural analysis
The most sophisticated systems analyse behaviour over time:
| Behaviour | Human | Bot |
|---|---|---|
| Mouse movement | Curved, varied, includes micro-movements | Absent, straight lines, or mathematically perfect curves |
| Scroll pattern | Variable speed, pauses | Absent or uniform |
| Click timing | Variable, includes hover-before-click | Instant, no hover |
| Page dwell time | Variable (reading time) | Instant or uniformly timed |
| Navigation pattern | Reads content, follows interest-based links | Sequential, exhaustive crawling |
Types of CAPTCHAs
reCAPTCHA v2 ("I'm not a robot" checkbox)
The checkbox CAPTCHA from Google. If the checkbox alone is insufficient (based on risk scoring), it presents an image challenge (select all images with traffic lights, etc.).
Risk score factors: IP reputation, cookies, browser history with Google services, mouse movement toward the checkbox.
reCAPTCHA v3 (invisible)
No visible challenge. Runs in the background and assigns a score (0.0 to 1.0) based on user behaviour. The website decides what score threshold triggers a block or additional verification.
How it works: A JavaScript snippet runs on the page, observes user behaviour and sends data to Google. The site receives a score and decides how to act.
hCaptcha
An alternative to reCAPTCHA that presents image-based challenges. Widely adopted after Google began charging for reCAPTCHA in some cases.
Cloudflare Turnstile
Cloudflare's managed challenge that replaces traditional CAPTCHAs. Uses a combination of browser challenges, proof-of-work and machine learning to verify visitors without visible challenges in most cases.
FunCaptcha (Arkose Labs)
Interactive challenges (rotate a 3D object, match an image) designed to be harder for automated solvers.
Prevention: Avoiding Detection
The most effective approach to bot detection is not solving CAPTCHAs, but avoiding triggering them in the first place.
Request-level best practices
Set complete headers:
headers = {
"User-Agent": (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/120.0.0.0 Safari/537.36"
),
"Accept": (
"text/html,application/xhtml+xml,application/xml;"
"q=0.9,image/webp,*/*;q=0.8"
),
"Accept-Language": "en-US,en;q=0.9",
"Accept-Encoding": "gzip, deflate, br",
"Connection": "keep-alive",
"Upgrade-Insecure-Requests": "1",
}
Maintain sessions: Use requests.Session() to persist cookies and connection state across requests.
Respect rate limits: Add random delays between requests (2-5 seconds minimum). Never hit the same domain more than once per second.
Browser-level best practices (headless browsers)
Patch navigator.webdriver:
// Playwright
await page.addInitScript(() => {
Object.defineProperty(navigator, "webdriver", {
get: () => undefined,
});
});
Set realistic viewport:
# Playwright
context = browser.new_context(
viewport={"width": 1920, "height": 1080},
screen={"width": 1920, "height": 1080},
)
Use stealth plugins:
- Puppeteer:
puppeteer-extra-plugin-stealthpatches multiple detection vectors. - Selenium:
undetected-chromedriverpatches ChromeDriver to avoid detection. - Playwright: No official stealth plugin, but the community maintains patches for common detection vectors.
Behavioural best practices
- Do not scrape every page on a site. Target only the pages that matter (contact pages, team pages).
- Mix up your request pattern. Do not visit pages in alphabetical or sequential order.
- Visit the homepage first before navigating to internal pages.
- Avoid scraping at perfectly regular intervals.
Handling CAPTCHAs When They Appear
Option 1: Slow down and retry
Many CAPTCHAs are triggered by rate limits. When you encounter one:
- Stop making requests to that domain.
- Wait 5-10 minutes.
- Clear cookies and use a different IP.
- Retry with longer delays between requests.
Option 2: Use a different approach
If a website consistently blocks scraping:
- Check for an API. Many websites with aggressive bot protection offer a public API for structured data access.
- Check for a data feed. Some sites publish RSS, sitemap or data export options.
- Use Email Extractor. If you can access the page in your browser, copy the page content and paste it into Email Extractor to extract email addresses without scraping.
- Use the Email Extractor browser extension. The browser extension extracts emails from pages you are already viewing in Chrome or Edge, without automated scraping.
Option 3: CAPTCHA solving services
Third-party services solve CAPTCHAs using human workers or machine learning:
| Service type | How it works | Typical speed |
|---|---|---|
| Human solving | Sends CAPTCHA image to a human worker who solves it | 10-30 seconds |
| AI solving | Machine learning models trained on specific CAPTCHA types | 1-10 seconds |
| Browser-based | Manages a browser session that handles challenges automatically | Variable |
Costs: Human solving typically costs $1-3 per 1,000 CAPTCHAs. AI solving costs less but has lower accuracy on complex challenges.
Ethical and legal considerations: Using CAPTCHA solving services circumvents a website's intentional access control. This may violate the website's terms of service and could raise legal issues under computer fraud and abuse laws in some jurisdictions. Evaluate the legal implications for your specific situation before using these services.
Bot Detection by Provider
Cloudflare
Detection level: Moderate to high (depends on the website's configuration).
How it works: Cloudflare sits between the visitor and the website. It evaluates each request based on IP reputation, browser fingerprint, challenge completion and behavioural signals.
Challenge types:
- JavaScript challenge (5-second page).
- Managed challenge (Turnstile).
- Interactive challenge (CAPTCHA when risk is high).
Common triggers:
- Datacenter IP addresses.
- Missing or unusual User-Agent.
- High request rate.
- Failed JavaScript challenge.
Akamai
Detection level: High.
How it works: Akamai's Bot Manager uses sensor data collected by client-side JavaScript, device fingerprinting and behavioural analysis.
Key detection: Akamai's sensor script generates a token that must be included in subsequent requests. Without executing the script and passing the token, requests are blocked.
PerimeterX (HUMAN)
Detection level: Very high.
How it works: Collects granular behavioural data (mouse movements, keystrokes, scroll patterns) and uses machine learning to distinguish human from automated behaviour. Specifically designed to defeat headless browsers.
When Not to Scrape
Sometimes the correct response to bot detection is to not scrape:
- The data is available through an API. Use the API instead.
- The data is available for purchase. Data providers may have the data you need at a reasonable cost.
- The website explicitly prohibits scraping. Respect terms of service, especially for sites where you have an account.
- The data is not public. Scraping data behind a login may violate computer access laws.
- The effort exceeds the value. If a site's bot protection requires extensive engineering to bypass, the cost may exceed the value of the data.
For extracting emails from files and documents you already have access to, use Email Extractor instead of scraping. Upload PDFs, spreadsheets, text files and other supported formats directly, and the tool extracts and deduplicates email addresses in your browser.