Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

How to Scrape Contact Pages at Scale Without Getting Blocked

On this page

The Problem

You have a list of company websites and you want to find contact email addresses from each one. Visiting each site manually takes minutes per company. At hundreds or thousands of targets, manual research is impractical.

The challenge is doing this at volume without getting blocked, missing data or violating the law.

Step 1: Find Contact Page URLs

Most websites put contact information at predictable URL paths. Before scraping the entire site, check the common paths first:

Common contact page paths

Path What it usually contains
/contact General contact form, sometimes email addresses
/contact-us Same as /contact
/about Company overview, sometimes team or leadership
/about-us Same as /about
/team Team members, sometimes with email addresses
/our-team Same as /team
/people Staff directory
/leadership Executive team
/staff Staff directory
/directory Staff or department directory
/support Support contact information
/press Media/PR contact with email
/media Same as /press
/careers Sometimes lists department heads or HR contacts
/imprint Legal contact (common in Germany/EU, required by law)
/impressum Same as /imprint (German)

Automated URL discovery

import requests

CONTACT_PATHS = [
    "/contact", "/contact-us", "/about", "/about-us",
    "/team", "/our-team", "/people", "/leadership",
    "/staff", "/directory", "/press", "/media",
    "/imprint", "/impressum", "/support",
]

def find_contact_pages(domain):
    """Try common contact page paths for a domain."""
    found = []
    base = f"https://{domain}"

    for path in CONTACT_PATHS:
        url = base + path
        try:
            resp = requests.head(
                url,
                timeout=5,
                allow_redirects=True,
                headers={"User-Agent": "Mozilla/5.0"}
            )
            if resp.status_code == 200:
                found.append(url)
        except requests.RequestException:
            continue

    return found

Using HEAD requests instead of GET is faster and lighter on the target server, since you only need to check whether the page exists.

Sitemap parsing

Many websites publish a sitemap at /sitemap.xml or list one in /robots.txt. Parsing the sitemap gives you every URL the site wants indexed, which you can filter for contact-related paths:

import xml.etree.ElementTree as ET

def find_contact_urls_in_sitemap(domain):
    """Parse sitemap.xml for contact-related URLs."""
    sitemap_url = f"https://{domain}/sitemap.xml"
    contact_keywords = [
        "contact", "team", "about", "people",
        "staff", "directory", "leadership", "press"
    ]

    try:
        resp = requests.get(
            sitemap_url, timeout=10,
            headers={"User-Agent": "Mozilla/5.0"}
        )
        if resp.status_code != 200:
            return []

        root = ET.fromstring(resp.content)
        # Handle XML namespace
        ns = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
        urls = [
            loc.text for loc in root.findall(".//sm:loc", ns)
        ]

        return [
            url for url in urls
            if any(kw in url.lower() for kw in contact_keywords)
        ]
    except Exception:
        return []

Step 2: Respect Rate Limits

Getting blocked is the primary obstacle to scraping at scale. Every website has limits on how many requests it will accept from a single source in a given time window.

Rate limiting strategies

Strategy Implementation Effect
Fixed delay time.sleep(2) between requests Simple but slow; 2 seconds per page means 30 pages/minute
Random delay time.sleep(random.uniform(1, 3)) Looks more natural than fixed intervals
Domain-level throttling Track per-domain request timing; enforce minimum interval per domain Allows fast overall throughput while being polite to each site
Adaptive delay Slow down when receiving 429 or 503 responses; speed up when responses are fast Adjusts to each server's capacity
Concurrent with limits Process multiple domains in parallel, but one request per domain at a time Maximises throughput without hammering any single site

Per-domain politeness

When scraping across thousands of different websites (one or two pages per site), the key rule is: never hit the same domain more than once every few seconds.

import time
from collections import defaultdict

last_request_time = defaultdict(float)
MIN_INTERVAL = 2.0  # seconds between requests to same domain

def rate_limited_get(url, domain):
    """Fetch URL with per-domain rate limiting."""
    elapsed = time.time() - last_request_time[domain]
    if elapsed < MIN_INTERVAL:
        time.sleep(MIN_INTERVAL - elapsed)

    response = requests.get(
        url, timeout=10,
        headers={"User-Agent": "Mozilla/5.0"}
    )
    last_request_time[domain] = time.time()
    return response

Robots.txt compliance

Before scraping any site, check its /robots.txt to see if the site explicitly disallows scraping of certain paths or by certain user agents:

from urllib.robotparser import RobotFileParser

def can_scrape(url, user_agent="*"):
    """Check if robots.txt allows scraping this URL."""
    from urllib.parse import urlparse
    parsed = urlparse(url)
    robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"

    rp = RobotFileParser()
    rp.set_url(robots_url)
    try:
        rp.read()
        return rp.can_fetch(user_agent, url)
    except Exception:
        return True  # If robots.txt is unreachable, proceed

Step 3: Handle Anti-Bot Measures

Common blocking methods

Method How it works How to handle
User-Agent filtering Blocks requests without a browser-like User-Agent Set a realistic User-Agent header
IP rate limiting Blocks IPs that send too many requests Rate limiting, proxy rotation
Cookie/session checks Requires cookies or a session token Use a session object that persists cookies
JavaScript challenge Serves a JS challenge page before the real content Use a headless browser (Playwright, Puppeteer)
CAPTCHA Requires human interaction Avoid triggering it (slower requests, realistic behaviour)
Cloudflare/Akamai CDN-level bot detection Headless browser with stealth patches
Honeypot links Invisible links that only bots click Only follow visible links

Session management

Using a requests Session persists cookies and connection state across requests, which looks more like normal browser behaviour:

session = requests.Session()
session.headers.update({
    "User-Agent": (
        "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
        "AppleWebKit/537.36 (KHTML, like Gecko) "
        "Chrome/120.0.0.0 Safari/537.36"
    ),
    "Accept": "text/html,application/xhtml+xml",
    "Accept-Language": "en-US,en;q=0.9",
})

Proxy rotation

For large-scale scraping (thousands of sites), rotating IP addresses prevents any single IP from being rate-limited:

Proxy type Cost Quality Speed
Datacenter proxies Low ($1-5/GB) Often detected as proxies Fast
Residential proxies Medium ($5-15/GB) Harder to detect Medium
Mobile proxies High ($15-30/GB) Very hard to detect Variable
ISP proxies Medium ($3-10/GB) Good balance of quality and cost Fast

For scraping company contact pages (not login-protected or heavily defended pages), datacenter proxies are usually sufficient.

When to use a headless browser

If more than 20-30% of your target sites serve JavaScript challenges or render contact information dynamically, switch from requests-based scraping to a headless browser:

from playwright.sync_api import sync_playwright

def scrape_with_browser(url):
    """Scrape a page that requires JavaScript rendering."""
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        context = browser.new_context(
            viewport={"width": 1920, "height": 1080},
            user_agent=(
                "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
                "AppleWebKit/537.36 Chrome/120.0.0.0 Safari/537.36"
            ),
        )
        page = context.new_page()
        page.goto(url, wait_until="networkidle", timeout=15000)
        content = page.content()
        browser.close()
        return content

The headless browser approach is slower (2-5 seconds per page vs 0.5-1 second for requests) but handles JavaScript-rendered content and is harder to detect as a bot.

Step 4: Extract Email Addresses

From HTML responses

import re
from bs4 import BeautifulSoup

def extract_emails_from_html(html):
    """Extract email addresses from HTML content."""
    # Decode HTML entities
    soup = BeautifulSoup(html, "html.parser")
    text = soup.get_text()

    # Also check href attributes for mailto: links
    mailto_emails = set()
    for link in soup.find_all("a", href=True):
        href = link["href"]
        if href.startswith("mailto:"):
            email = href.replace("mailto:", "").split("?")[0].strip()
            mailto_emails.add(email.lower())

    # Regex extraction from page text
    pattern = r"[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}"
    text_emails = {e.lower() for e in re.findall(pattern, text)}

    # Also search raw HTML for obfuscated emails
    raw_emails = {e.lower() for e in re.findall(pattern, html)}

    return mailto_emails | text_emails | raw_emails

From downloaded files

If contact pages link to downloadable files (PDF directories, vCards, spreadsheets), download these files and upload them to Email Extractor for extraction. The tool supports PDF, XLSX, DOCX, CSV, VCF and other formats, and handles deduplication automatically.

Step 5: Process and Deduplicate

After scraping across many websites, you will have email addresses from hundreds of pages. Processing steps:

Clean and normalise

def clean_email(email):
    """Normalise an extracted email address."""
    email = email.lower().strip()

    # Remove common false positive suffixes
    for suffix in [".png", ".jpg", ".gif", ".css", ".js"]:
        if email.endswith(suffix):
            return None

    # Basic validation
    if "@" not in email or "." not in email.split("@")[1]:
        return None

    return email

Deduplicate across sources

Upload your collected results to Email Extractor:

  1. Save all extracted emails to a text file or CSV (one email per line).
  2. Go to Email Extractor.
  3. Select "Text and files."
  4. Upload the file (or paste the text).
  5. Click "Extract emails."
  6. Download the deduplicated result as CSV.

The tool performs case-insensitive deduplication, so User@Example.com and user@example.com are treated as one address.

Associate emails with sources

Track which website each email came from. This is important for personalisation and compliance:

results = []
for domain, emails in domain_emails.items():
    for email in emails:
        results.append({
            "email": email,
            "source_domain": domain,
            "source_url": url,
            "found_date": datetime.date.today().isoformat(),
        })

Scraping public contact pages is a common business practice, but legal constraints apply:

Consideration Guidance
robots.txt Respect disallow directives. While not legally binding in all jurisdictions, violating robots.txt weakens your legal position.
Terms of service Some websites prohibit scraping in their terms. Scraping in violation of terms may create legal risk, especially after the LinkedIn v. hiQ case clarified the boundary for publicly accessible data.
GDPR (EU/UK) Collecting personal data (email addresses) from public sources requires a lawful basis. Legitimate interest may apply for B2B contacts, but you must be able to demonstrate the interest and respond to data subject rights requests.
CAN-SPAM (US) Does not restrict collection, but regulates how collected addresses can be used for commercial email.
CCPA (California) Gives individuals rights over personal information, including data collected from public sources. Must honour opt-out requests.
CASL (Canada) Requires consent before sending commercial electronic messages. Collecting addresses does not grant consent.

Bottom line: Collecting publicly available business email addresses from company websites is generally permissible. How you use those addresses is where regulation gets strict. Consult legal counsel for your specific jurisdiction and use case.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)