Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

How to Build a Simple Email Scraper with Python (Beginner Tutorial)

On this page

What You Will Build

This tutorial walks through building a Python script that:

  1. Fetches a web page.
  2. Extracts all email addresses from the page content.
  3. Follows internal links to find emails on related pages.
  4. Deduplicates and saves the results.

This is a learning exercise. For extracting emails from files you already have (PDFs, spreadsheets, text files, HTML files), Email Extractor handles this without writing code. The tool supports 19 file types, deduplicates automatically and runs entirely in your browser.

Prerequisites

Python 3.8 or later. Check your version:

python3 --version

Install required libraries:

pip install requests beautifulsoup4
  • requests fetches web pages.
  • beautifulsoup4 parses HTML.
  • re (regex) is built into Python and extracts email patterns from text.

Step 1: Fetch a Web Page

import requests

def fetch_page(url):
    """Fetch a web page and return its HTML content."""
    headers = {
        "User-Agent": (
            "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
            "AppleWebKit/537.36 (KHTML, like Gecko) "
            "Chrome/120.0.0.0 Safari/537.36"
        )
    }
    try:
        response = requests.get(url, headers=headers, timeout=10)
        response.raise_for_status()
        return response.text
    except requests.RequestException as e:
        print(f"Failed to fetch {url}: {e}")
        return None

What this does:

  • Sets a User-Agent header so the request looks like a normal browser visit. Some websites block requests without a User-Agent.
  • Sets a 10-second timeout to avoid hanging on unresponsive servers.
  • Returns None if the request fails, so the rest of the script can handle the failure gracefully.

Step 2: Extract Email Addresses with Regex

import re

def extract_emails(html):
    """Extract email addresses from HTML content."""
    # Standard email regex pattern
    pattern = r"[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}"
    
    emails = set(re.findall(pattern, html))
    
    # Normalise to lowercase for deduplication
    emails = {email.lower() for email in emails}
    
    return emails

About the regex pattern:

Part Matches
[a-zA-Z0-9._%+\-]+ Local part (before @): letters, digits, dots, underscores, percent, plus, hyphens
@ The @ symbol
[a-zA-Z0-9.\-]+ Domain name: letters, digits, dots, hyphens
\.[a-zA-Z]{2,} Top-level domain: dot followed by 2+ letters

Limitations of this pattern:

  • Does not match internationalised email addresses (IDN domains, non-ASCII local parts).
  • May match text that looks like an email but is not one (e.g., version numbers like v2.0@release).
  • Does not validate that the address is deliverable. It only finds strings that look like email addresses.

To find emails across multiple pages on a site, extract internal links from each page:

from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse

def extract_links(html, base_url):
    """Extract internal links from an HTML page."""
    soup = BeautifulSoup(html, "html.parser")
    links = set()
    
    base_domain = urlparse(base_url).netloc
    
    for anchor in soup.find_all("a", href=True):
        href = anchor["href"]
        
        # Resolve relative URLs
        full_url = urljoin(base_url, href)
        
        # Only follow links on the same domain
        if urlparse(full_url).netloc == base_domain:
            # Remove fragments (#section)
            full_url = full_url.split("#")[0]
            # Skip non-HTTP links
            if full_url.startswith(("http://", "https://")):
                links.add(full_url)
    
    return links

What this does:

  • Parses the HTML with BeautifulSoup to find all <a> tags with href attributes.
  • Uses urljoin to handle relative URLs (e.g., /contact becomes https://example.com/contact).
  • Filters to only same-domain links (internal links), so the scraper does not wander to other websites.
  • Removes URL fragments (#section) to avoid visiting the same page multiple times.

Step 4: Build the Crawler

Combine the pieces into a crawler that visits multiple pages:

import time

def crawl(start_url, max_pages=50):
    """Crawl a website and extract all email addresses."""
    visited = set()
    to_visit = {start_url}
    all_emails = set()
    
    while to_visit and len(visited) < max_pages:
        url = to_visit.pop()
        
        if url in visited:
            continue
        
        print(f"Visiting: {url}")
        visited.add(url)
        
        html = fetch_page(url)
        if html is None:
            continue
        
        # Extract emails from this page
        emails = extract_emails(html)
        if emails:
            print(f"  Found {len(emails)} email(s)")
            all_emails.update(emails)
        
        # Find more pages to visit
        links = extract_links(html, url)
        new_links = links - visited
        to_visit.update(new_links)
        
        # Rate limiting: wait between requests
        time.sleep(2)
    
    return all_emails

Key design decisions:

  • max_pages limit. Prevents the scraper from running indefinitely on large sites.
  • Visited tracking. Avoids processing the same page twice.
  • Rate limiting. A 2-second delay between requests is a minimum courtesy. Adjust based on the site's terms and your relationship with the site owner.
  • Breadth-first crawling. to_visit is a set (unordered), so this is roughly breadth-first. For depth-first, use a list and pop() from the end.

Step 5: Save Results

import csv

def save_emails(emails, filename="extracted_emails.csv"):
    """Save extracted emails to a CSV file."""
    sorted_emails = sorted(emails)
    
    with open(filename, "w", newline="") as f:
        writer = csv.writer(f)
        writer.writerow(["email"])
        for email in sorted_emails:
            writer.writerow([email])
    
    print(f"Saved {len(sorted_emails)} unique emails to {filename}")

Complete Script

import re
import csv
import time
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse


def fetch_page(url):
    """Fetch a web page and return its HTML content."""
    headers = {
        "User-Agent": (
            "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
            "AppleWebKit/537.36 (KHTML, like Gecko) "
            "Chrome/120.0.0.0 Safari/537.36"
        )
    }
    try:
        response = requests.get(url, headers=headers, timeout=10)
        response.raise_for_status()
        return response.text
    except requests.RequestException as e:
        print(f"Failed to fetch {url}: {e}")
        return None


def extract_emails(html):
    """Extract email addresses from HTML content."""
    pattern = r"[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}"
    emails = set(re.findall(pattern, html))
    emails = {email.lower() for email in emails}
    return emails


def extract_links(html, base_url):
    """Extract internal links from an HTML page."""
    soup = BeautifulSoup(html, "html.parser")
    links = set()
    base_domain = urlparse(base_url).netloc

    for anchor in soup.find_all("a", href=True):
        href = anchor["href"]
        full_url = urljoin(base_url, href)
        if urlparse(full_url).netloc == base_domain:
            full_url = full_url.split("#")[0]
            if full_url.startswith(("http://", "https://")):
                links.add(full_url)

    return links


def crawl(start_url, max_pages=50):
    """Crawl a website and extract all email addresses."""
    visited = set()
    to_visit = {start_url}
    all_emails = set()

    while to_visit and len(visited) < max_pages:
        url = to_visit.pop()
        if url in visited:
            continue

        print(f"Visiting: {url}")
        visited.add(url)

        html = fetch_page(url)
        if html is None:
            continue

        emails = extract_emails(html)
        if emails:
            print(f"  Found {len(emails)} email(s)")
            all_emails.update(emails)

        links = extract_links(html, url)
        to_visit.update(links - visited)

        time.sleep(2)

    return all_emails


def save_emails(emails, filename="extracted_emails.csv"):
    """Save extracted emails to a CSV file."""
    sorted_emails = sorted(emails)
    with open(filename, "w", newline="") as f:
        writer = csv.writer(f)
        writer.writerow(["email"])
        for email in sorted_emails:
            writer.writerow([email])
    print(f"Saved {len(sorted_emails)} unique emails to {filename}")


if __name__ == "__main__":
    target_url = "https://example.com"
    emails = crawl(target_url, max_pages=50)
    save_emails(emails)

Running the Script

python3 scraper.py

Expected output:

Visiting: https://example.com
Visiting: https://example.com/about
Visiting: https://example.com/contact
  Found 3 email(s)
Visiting: https://example.com/team
  Found 5 email(s)
Saved 7 unique emails to extracted_emails.csv

Handling Common Issues

JavaScript-rendered content

This script only works with content in the initial HTML response. If the website loads email addresses with JavaScript after the page loads, you need a headless browser instead:

from playwright.sync_api import sync_playwright

def fetch_page_with_js(url):
    """Fetch a page that requires JavaScript rendering."""
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto(url, wait_until="networkidle")
        content = page.content()
        browser.close()
        return content

Install Playwright with pip install playwright && playwright install chromium.

Obfuscated email addresses

Some websites obfuscate email addresses to prevent scraping:

Obfuscation method Example Solution
HTML entities user&#64;example.com BeautifulSoup decodes these automatically
[at] substitution user [at] example.com Add regex pattern: [a-zA-Z0-9._%+-]+\s*\[at\]\s*[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}
JavaScript assembly Build email from variables in JS Requires headless browser to render
Image-based Email displayed as an image Requires OCR (not covered here)
Contact form only No email visible Cannot be scraped

Rate limiting and blocking

If the website blocks your scraper:

  1. Increase the delay between requests. Try 5-10 seconds.
  2. Check robots.txt. Visit https://example.com/robots.txt to see if the site disallows scraping specific paths.
  3. Respect the response. If you receive a 429 (Too Many Requests) or 403 (Forbidden) response, stop.

Filtering false positives

The regex may match strings that look like email addresses but are not:

def filter_emails(emails):
    """Remove common false positives."""
    excluded_extensions = {
        ".png", ".jpg", ".jpeg", ".gif", ".svg",
        ".css", ".js", ".woff", ".woff2", ".ttf"
    }
    
    filtered = set()
    for email in emails:
        # Skip image and asset filenames
        if any(email.endswith(ext) for ext in excluded_extensions):
            continue
        # Skip very long local parts (likely not real)
        local_part = email.split("@")[0]
        if len(local_part) > 64:
            continue
        filtered.add(email)
    
    return filtered

Limitations of This Approach

Static HTML only. Does not handle JavaScript-rendered pages without a headless browser.

No email verification. The script finds text that matches email patterns. It does not check whether the addresses are active or deliverable.

No consent. Finding an email address on a public web page does not grant permission to send marketing email. Review CAN-SPAM, GDPR and CASL requirements before using scraped addresses for outreach.

No contact data. The script extracts email addresses only, not names, titles or company information. To associate emails with names, you would need additional parsing logic specific to each site's HTML structure.

Single-threaded. The script visits one page at a time. For faster scraping, consider using asyncio with aiohttp, or multi-threading with concurrent.futures.

When to Use This vs Email Extractor

Scenario Use this Python script Use Email Extractor
Scrape emails from live web pages Yes No (use webpage extraction for public URLs)
Extract emails from downloaded files No Yes (supports 19 file types)
Extract from pasted text No Yes
Need deduplication Write it yourself Built-in
Need source tracking Write it yourself Built-in (CSV with sources)
No coding required No Yes

For most non-technical users, Email Extractor handles extraction from files and text without any coding. The Python approach is useful when you need to scrape live websites, customise the extraction logic or integrate scraping into a larger automated pipeline.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)