How to Scrape Contact Pages at Scale Without Getting Blocked
On this page
The Problem
You have a list of company websites and you want to find contact email addresses from each one. Visiting each site manually takes minutes per company. At hundreds or thousands of targets, manual research is impractical.
The challenge is doing this at volume without getting blocked, missing data or violating the law.
Step 1: Find Contact Page URLs
Most websites put contact information at predictable URL paths. Before scraping the entire site, check the common paths first:
Common contact page paths
| Path | What it usually contains |
|---|---|
| /contact | General contact form, sometimes email addresses |
| /contact-us | Same as /contact |
| /about | Company overview, sometimes team or leadership |
| /about-us | Same as /about |
| /team | Team members, sometimes with email addresses |
| /our-team | Same as /team |
| /people | Staff directory |
| /leadership | Executive team |
| /staff | Staff directory |
| /directory | Staff or department directory |
| /support | Support contact information |
| /press | Media/PR contact with email |
| /media | Same as /press |
| /careers | Sometimes lists department heads or HR contacts |
| /imprint | Legal contact (common in Germany/EU, required by law) |
| /impressum | Same as /imprint (German) |
Automated URL discovery
import requests
CONTACT_PATHS = [
"/contact", "/contact-us", "/about", "/about-us",
"/team", "/our-team", "/people", "/leadership",
"/staff", "/directory", "/press", "/media",
"/imprint", "/impressum", "/support",
]
def find_contact_pages(domain):
"""Try common contact page paths for a domain."""
found = []
base = f"https://{domain}"
for path in CONTACT_PATHS:
url = base + path
try:
resp = requests.head(
url,
timeout=5,
allow_redirects=True,
headers={"User-Agent": "Mozilla/5.0"}
)
if resp.status_code == 200:
found.append(url)
except requests.RequestException:
continue
return found
Using HEAD requests instead of GET is faster and lighter on the target server, since you only need to check whether the page exists.
Sitemap parsing
Many websites publish a sitemap at /sitemap.xml or list one in /robots.txt. Parsing the sitemap gives you every URL the site wants indexed, which you can filter for contact-related paths:
import xml.etree.ElementTree as ET
def find_contact_urls_in_sitemap(domain):
"""Parse sitemap.xml for contact-related URLs."""
sitemap_url = f"https://{domain}/sitemap.xml"
contact_keywords = [
"contact", "team", "about", "people",
"staff", "directory", "leadership", "press"
]
try:
resp = requests.get(
sitemap_url, timeout=10,
headers={"User-Agent": "Mozilla/5.0"}
)
if resp.status_code != 200:
return []
root = ET.fromstring(resp.content)
# Handle XML namespace
ns = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
urls = [
loc.text for loc in root.findall(".//sm:loc", ns)
]
return [
url for url in urls
if any(kw in url.lower() for kw in contact_keywords)
]
except Exception:
return []
Step 2: Respect Rate Limits
Getting blocked is the primary obstacle to scraping at scale. Every website has limits on how many requests it will accept from a single source in a given time window.
Rate limiting strategies
| Strategy | Implementation | Effect |
|---|---|---|
| Fixed delay | time.sleep(2) between requests |
Simple but slow; 2 seconds per page means 30 pages/minute |
| Random delay | time.sleep(random.uniform(1, 3)) |
Looks more natural than fixed intervals |
| Domain-level throttling | Track per-domain request timing; enforce minimum interval per domain | Allows fast overall throughput while being polite to each site |
| Adaptive delay | Slow down when receiving 429 or 503 responses; speed up when responses are fast | Adjusts to each server's capacity |
| Concurrent with limits | Process multiple domains in parallel, but one request per domain at a time | Maximises throughput without hammering any single site |
Per-domain politeness
When scraping across thousands of different websites (one or two pages per site), the key rule is: never hit the same domain more than once every few seconds.
import time
from collections import defaultdict
last_request_time = defaultdict(float)
MIN_INTERVAL = 2.0 # seconds between requests to same domain
def rate_limited_get(url, domain):
"""Fetch URL with per-domain rate limiting."""
elapsed = time.time() - last_request_time[domain]
if elapsed < MIN_INTERVAL:
time.sleep(MIN_INTERVAL - elapsed)
response = requests.get(
url, timeout=10,
headers={"User-Agent": "Mozilla/5.0"}
)
last_request_time[domain] = time.time()
return response
Robots.txt compliance
Before scraping any site, check its /robots.txt to see if the site explicitly disallows scraping of certain paths or by certain user agents:
from urllib.robotparser import RobotFileParser
def can_scrape(url, user_agent="*"):
"""Check if robots.txt allows scraping this URL."""
from urllib.parse import urlparse
parsed = urlparse(url)
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
rp = RobotFileParser()
rp.set_url(robots_url)
try:
rp.read()
return rp.can_fetch(user_agent, url)
except Exception:
return True # If robots.txt is unreachable, proceed
Step 3: Handle Anti-Bot Measures
Common blocking methods
| Method | How it works | How to handle |
|---|---|---|
| User-Agent filtering | Blocks requests without a browser-like User-Agent | Set a realistic User-Agent header |
| IP rate limiting | Blocks IPs that send too many requests | Rate limiting, proxy rotation |
| Cookie/session checks | Requires cookies or a session token | Use a session object that persists cookies |
| JavaScript challenge | Serves a JS challenge page before the real content | Use a headless browser (Playwright, Puppeteer) |
| CAPTCHA | Requires human interaction | Avoid triggering it (slower requests, realistic behaviour) |
| Cloudflare/Akamai | CDN-level bot detection | Headless browser with stealth patches |
| Honeypot links | Invisible links that only bots click | Only follow visible links |
Session management
Using a requests Session persists cookies and connection state across requests, which looks more like normal browser behaviour:
session = requests.Session()
session.headers.update({
"User-Agent": (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/120.0.0.0 Safari/537.36"
),
"Accept": "text/html,application/xhtml+xml",
"Accept-Language": "en-US,en;q=0.9",
})
Proxy rotation
For large-scale scraping (thousands of sites), rotating IP addresses prevents any single IP from being rate-limited:
| Proxy type | Cost | Quality | Speed |
|---|---|---|---|
| Datacenter proxies | Low ($1-5/GB) | Often detected as proxies | Fast |
| Residential proxies | Medium ($5-15/GB) | Harder to detect | Medium |
| Mobile proxies | High ($15-30/GB) | Very hard to detect | Variable |
| ISP proxies | Medium ($3-10/GB) | Good balance of quality and cost | Fast |
For scraping company contact pages (not login-protected or heavily defended pages), datacenter proxies are usually sufficient.
When to use a headless browser
If more than 20-30% of your target sites serve JavaScript challenges or render contact information dynamically, switch from requests-based scraping to a headless browser:
from playwright.sync_api import sync_playwright
def scrape_with_browser(url):
"""Scrape a page that requires JavaScript rendering."""
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
context = browser.new_context(
viewport={"width": 1920, "height": 1080},
user_agent=(
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 Chrome/120.0.0.0 Safari/537.36"
),
)
page = context.new_page()
page.goto(url, wait_until="networkidle", timeout=15000)
content = page.content()
browser.close()
return content
The headless browser approach is slower (2-5 seconds per page vs 0.5-1 second for requests) but handles JavaScript-rendered content and is harder to detect as a bot.
Step 4: Extract Email Addresses
From HTML responses
import re
from bs4 import BeautifulSoup
def extract_emails_from_html(html):
"""Extract email addresses from HTML content."""
# Decode HTML entities
soup = BeautifulSoup(html, "html.parser")
text = soup.get_text()
# Also check href attributes for mailto: links
mailto_emails = set()
for link in soup.find_all("a", href=True):
href = link["href"]
if href.startswith("mailto:"):
email = href.replace("mailto:", "").split("?")[0].strip()
mailto_emails.add(email.lower())
# Regex extraction from page text
pattern = r"[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}"
text_emails = {e.lower() for e in re.findall(pattern, text)}
# Also search raw HTML for obfuscated emails
raw_emails = {e.lower() for e in re.findall(pattern, html)}
return mailto_emails | text_emails | raw_emails
From downloaded files
If contact pages link to downloadable files (PDF directories, vCards, spreadsheets), download these files and upload them to Email Extractor for extraction. The tool supports PDF, XLSX, DOCX, CSV, VCF and other formats, and handles deduplication automatically.
Step 5: Process and Deduplicate
After scraping across many websites, you will have email addresses from hundreds of pages. Processing steps:
Clean and normalise
def clean_email(email):
"""Normalise an extracted email address."""
email = email.lower().strip()
# Remove common false positive suffixes
for suffix in [".png", ".jpg", ".gif", ".css", ".js"]:
if email.endswith(suffix):
return None
# Basic validation
if "@" not in email or "." not in email.split("@")[1]:
return None
return email
Deduplicate across sources
Upload your collected results to Email Extractor:
- Save all extracted emails to a text file or CSV (one email per line).
- Go to Email Extractor.
- Select "Text and files."
- Upload the file (or paste the text).
- Click "Extract emails."
- Download the deduplicated result as CSV.
The tool performs case-insensitive deduplication, so User@Example.com and user@example.com are treated as one address.
Associate emails with sources
Track which website each email came from. This is important for personalisation and compliance:
results = []
for domain, emails in domain_emails.items():
for email in emails:
results.append({
"email": email,
"source_domain": domain,
"source_url": url,
"found_date": datetime.date.today().isoformat(),
})
Legal Considerations
Scraping public contact pages is a common business practice, but legal constraints apply:
| Consideration | Guidance |
|---|---|
| robots.txt | Respect disallow directives. While not legally binding in all jurisdictions, violating robots.txt weakens your legal position. |
| Terms of service | Some websites prohibit scraping in their terms. Scraping in violation of terms may create legal risk, especially after the LinkedIn v. hiQ case clarified the boundary for publicly accessible data. |
| GDPR (EU/UK) | Collecting personal data (email addresses) from public sources requires a lawful basis. Legitimate interest may apply for B2B contacts, but you must be able to demonstrate the interest and respond to data subject rights requests. |
| CAN-SPAM (US) | Does not restrict collection, but regulates how collected addresses can be used for commercial email. |
| CCPA (California) | Gives individuals rights over personal information, including data collected from public sources. Must honour opt-out requests. |
| CASL (Canada) | Requires consent before sending commercial electronic messages. Collecting addresses does not grant consent. |
Bottom line: Collecting publicly available business email addresses from company websites is generally permissible. How you use those addresses is where regulation gets strict. Consult legal counsel for your specific jurisdiction and use case.