How to Build a Simple Email Scraper with Python (Beginner Tutorial)
On this page
What You Will Build
This tutorial walks through building a Python script that:
- Fetches a web page.
- Extracts all email addresses from the page content.
- Follows internal links to find emails on related pages.
- Deduplicates and saves the results.
This is a learning exercise. For extracting emails from files you already have (PDFs, spreadsheets, text files, HTML files), Email Extractor handles this without writing code. The tool supports 19 file types, deduplicates automatically and runs entirely in your browser.
Prerequisites
Python 3.8 or later. Check your version:
python3 --version
Install required libraries:
pip install requests beautifulsoup4
requestsfetches web pages.beautifulsoup4parses HTML.re(regex) is built into Python and extracts email patterns from text.
Step 1: Fetch a Web Page
import requests
def fetch_page(url):
"""Fetch a web page and return its HTML content."""
headers = {
"User-Agent": (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/120.0.0.0 Safari/537.36"
)
}
try:
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
return response.text
except requests.RequestException as e:
print(f"Failed to fetch {url}: {e}")
return None
What this does:
- Sets a User-Agent header so the request looks like a normal browser visit. Some websites block requests without a User-Agent.
- Sets a 10-second timeout to avoid hanging on unresponsive servers.
- Returns None if the request fails, so the rest of the script can handle the failure gracefully.
Step 2: Extract Email Addresses with Regex
import re
def extract_emails(html):
"""Extract email addresses from HTML content."""
# Standard email regex pattern
pattern = r"[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}"
emails = set(re.findall(pattern, html))
# Normalise to lowercase for deduplication
emails = {email.lower() for email in emails}
return emails
About the regex pattern:
| Part | Matches |
|---|---|
[a-zA-Z0-9._%+\-]+ |
Local part (before @): letters, digits, dots, underscores, percent, plus, hyphens |
@ |
The @ symbol |
[a-zA-Z0-9.\-]+ |
Domain name: letters, digits, dots, hyphens |
\.[a-zA-Z]{2,} |
Top-level domain: dot followed by 2+ letters |
Limitations of this pattern:
- Does not match internationalised email addresses (IDN domains, non-ASCII local parts).
- May match text that looks like an email but is not one (e.g., version numbers like
v2.0@release). - Does not validate that the address is deliverable. It only finds strings that look like email addresses.
Step 3: Extract Links for Crawling
To find emails across multiple pages on a site, extract internal links from each page:
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse
def extract_links(html, base_url):
"""Extract internal links from an HTML page."""
soup = BeautifulSoup(html, "html.parser")
links = set()
base_domain = urlparse(base_url).netloc
for anchor in soup.find_all("a", href=True):
href = anchor["href"]
# Resolve relative URLs
full_url = urljoin(base_url, href)
# Only follow links on the same domain
if urlparse(full_url).netloc == base_domain:
# Remove fragments (#section)
full_url = full_url.split("#")[0]
# Skip non-HTTP links
if full_url.startswith(("http://", "https://")):
links.add(full_url)
return links
What this does:
- Parses the HTML with BeautifulSoup to find all
<a>tags withhrefattributes. - Uses
urljointo handle relative URLs (e.g.,/contactbecomeshttps://example.com/contact). - Filters to only same-domain links (internal links), so the scraper does not wander to other websites.
- Removes URL fragments (#section) to avoid visiting the same page multiple times.
Step 4: Build the Crawler
Combine the pieces into a crawler that visits multiple pages:
import time
def crawl(start_url, max_pages=50):
"""Crawl a website and extract all email addresses."""
visited = set()
to_visit = {start_url}
all_emails = set()
while to_visit and len(visited) < max_pages:
url = to_visit.pop()
if url in visited:
continue
print(f"Visiting: {url}")
visited.add(url)
html = fetch_page(url)
if html is None:
continue
# Extract emails from this page
emails = extract_emails(html)
if emails:
print(f" Found {len(emails)} email(s)")
all_emails.update(emails)
# Find more pages to visit
links = extract_links(html, url)
new_links = links - visited
to_visit.update(new_links)
# Rate limiting: wait between requests
time.sleep(2)
return all_emails
Key design decisions:
- max_pages limit. Prevents the scraper from running indefinitely on large sites.
- Visited tracking. Avoids processing the same page twice.
- Rate limiting. A 2-second delay between requests is a minimum courtesy. Adjust based on the site's terms and your relationship with the site owner.
- Breadth-first crawling.
to_visitis a set (unordered), so this is roughly breadth-first. For depth-first, use a list andpop()from the end.
Step 5: Save Results
import csv
def save_emails(emails, filename="extracted_emails.csv"):
"""Save extracted emails to a CSV file."""
sorted_emails = sorted(emails)
with open(filename, "w", newline="") as f:
writer = csv.writer(f)
writer.writerow(["email"])
for email in sorted_emails:
writer.writerow([email])
print(f"Saved {len(sorted_emails)} unique emails to {filename}")
Complete Script
import re
import csv
import time
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse
def fetch_page(url):
"""Fetch a web page and return its HTML content."""
headers = {
"User-Agent": (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/120.0.0.0 Safari/537.36"
)
}
try:
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
return response.text
except requests.RequestException as e:
print(f"Failed to fetch {url}: {e}")
return None
def extract_emails(html):
"""Extract email addresses from HTML content."""
pattern = r"[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}"
emails = set(re.findall(pattern, html))
emails = {email.lower() for email in emails}
return emails
def extract_links(html, base_url):
"""Extract internal links from an HTML page."""
soup = BeautifulSoup(html, "html.parser")
links = set()
base_domain = urlparse(base_url).netloc
for anchor in soup.find_all("a", href=True):
href = anchor["href"]
full_url = urljoin(base_url, href)
if urlparse(full_url).netloc == base_domain:
full_url = full_url.split("#")[0]
if full_url.startswith(("http://", "https://")):
links.add(full_url)
return links
def crawl(start_url, max_pages=50):
"""Crawl a website and extract all email addresses."""
visited = set()
to_visit = {start_url}
all_emails = set()
while to_visit and len(visited) < max_pages:
url = to_visit.pop()
if url in visited:
continue
print(f"Visiting: {url}")
visited.add(url)
html = fetch_page(url)
if html is None:
continue
emails = extract_emails(html)
if emails:
print(f" Found {len(emails)} email(s)")
all_emails.update(emails)
links = extract_links(html, url)
to_visit.update(links - visited)
time.sleep(2)
return all_emails
def save_emails(emails, filename="extracted_emails.csv"):
"""Save extracted emails to a CSV file."""
sorted_emails = sorted(emails)
with open(filename, "w", newline="") as f:
writer = csv.writer(f)
writer.writerow(["email"])
for email in sorted_emails:
writer.writerow([email])
print(f"Saved {len(sorted_emails)} unique emails to {filename}")
if __name__ == "__main__":
target_url = "https://example.com"
emails = crawl(target_url, max_pages=50)
save_emails(emails)
Running the Script
python3 scraper.py
Expected output:
Visiting: https://example.com
Visiting: https://example.com/about
Visiting: https://example.com/contact
Found 3 email(s)
Visiting: https://example.com/team
Found 5 email(s)
Saved 7 unique emails to extracted_emails.csv
Handling Common Issues
JavaScript-rendered content
This script only works with content in the initial HTML response. If the website loads email addresses with JavaScript after the page loads, you need a headless browser instead:
from playwright.sync_api import sync_playwright
def fetch_page_with_js(url):
"""Fetch a page that requires JavaScript rendering."""
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto(url, wait_until="networkidle")
content = page.content()
browser.close()
return content
Install Playwright with pip install playwright && playwright install chromium.
Obfuscated email addresses
Some websites obfuscate email addresses to prevent scraping:
| Obfuscation method | Example | Solution |
|---|---|---|
| HTML entities | user@example.com |
BeautifulSoup decodes these automatically |
| [at] substitution | user [at] example.com |
Add regex pattern: [a-zA-Z0-9._%+-]+\s*\[at\]\s*[a-zA-Z0-9.-]+\.[a-zA-Z]{2,} |
| JavaScript assembly | Build email from variables in JS | Requires headless browser to render |
| Image-based | Email displayed as an image | Requires OCR (not covered here) |
| Contact form only | No email visible | Cannot be scraped |
Rate limiting and blocking
If the website blocks your scraper:
- Increase the delay between requests. Try 5-10 seconds.
- Check robots.txt. Visit
https://example.com/robots.txtto see if the site disallows scraping specific paths. - Respect the response. If you receive a 429 (Too Many Requests) or 403 (Forbidden) response, stop.
Filtering false positives
The regex may match strings that look like email addresses but are not:
def filter_emails(emails):
"""Remove common false positives."""
excluded_extensions = {
".png", ".jpg", ".jpeg", ".gif", ".svg",
".css", ".js", ".woff", ".woff2", ".ttf"
}
filtered = set()
for email in emails:
# Skip image and asset filenames
if any(email.endswith(ext) for ext in excluded_extensions):
continue
# Skip very long local parts (likely not real)
local_part = email.split("@")[0]
if len(local_part) > 64:
continue
filtered.add(email)
return filtered
Limitations of This Approach
Static HTML only. Does not handle JavaScript-rendered pages without a headless browser.
No email verification. The script finds text that matches email patterns. It does not check whether the addresses are active or deliverable.
No consent. Finding an email address on a public web page does not grant permission to send marketing email. Review CAN-SPAM, GDPR and CASL requirements before using scraped addresses for outreach.
No contact data. The script extracts email addresses only, not names, titles or company information. To associate emails with names, you would need additional parsing logic specific to each site's HTML structure.
Single-threaded. The script visits one page at a time. For faster scraping, consider using asyncio with aiohttp, or multi-threading with concurrent.futures.
When to Use This vs Email Extractor
| Scenario | Use this Python script | Use Email Extractor |
|---|---|---|
| Scrape emails from live web pages | Yes | No (use webpage extraction for public URLs) |
| Extract emails from downloaded files | No | Yes (supports 19 file types) |
| Extract from pasted text | No | Yes |
| Need deduplication | Write it yourself | Built-in |
| Need source tracking | Write it yourself | Built-in (CSV with sources) |
| No coding required | No | Yes |
For most non-technical users, Email Extractor handles extraction from files and text without any coding. The Python approach is useful when you need to scrape live websites, customise the extraction logic or integrate scraping into a larger automated pipeline.
Related Guides
- How to Scrape Emails from Any Website (Ethically)
- Regex for Email Extraction
- Headless Browsers for Scraping: Puppeteer vs Playwright vs Selenium
- Web Scraping vs Web Crawling
- Scraping Ethics and Best Practices