Web Scraping for Lead Generation: How to Build Prospect Lists from Public Data
On this page
What Scraping Adds to Lead Generation
B2B data providers (Apollo, ZoomInfo, Lusha) give you large databases of contacts to search. Web scraping gives you data they may not have: contacts from niche directories, industry-specific listings, conference speaker pages, company team pages and public records that are not covered by the major providers.
Scraping is not a replacement for data providers. It is a supplement for specific use cases where the data you need exists publicly on the web but is not aggregated into a searchable database.
Where to Scrape for Leads
Company team pages
Many companies list their team on an "About Us" or "Team" page, often including names, titles and sometimes email addresses.
What you find: Names, job titles, photos, LinkedIn profiles. Email addresses are less common but sometimes included.
Use case: Identifying decision-makers at target companies when data providers lack coverage (small companies, startups, international firms).
How: Scrape the team page for names and titles. Use an email finder tool to look up email addresses by name and company domain.
Industry directories
Trade associations, professional organisations and industry bodies maintain member directories.
Examples:
- Chamber of Commerce directories.
- Legal firm directories (Martindale-Hubbell, Avvo).
- Medical provider directories (NPI registry, state licensing boards).
- Real estate agent directories (Realtor.com, state association sites).
- Technology company directories (Clutch, G2, BuiltWith).
- Accounting firm directories (AICPA, state CPA societies).
What you find: Company name, contact person, email, phone, address, specialisation.
Use case: Building targeted lists of businesses in a specific industry and geography.
Conference and event speaker pages
Conferences publish speaker bios, company affiliations and sometimes contact information.
What you find: Names, titles, company names, headshot photos, bio text, social links.
Use case: Building lists of thought leaders and decision-makers in your target market. Conference speakers are typically senior professionals who are active in their field.
Job postings
Job boards reveal which companies are hiring for specific roles, which signals their priorities and growth areas.
What you find: Company name, job title, location, requirements, sometimes a hiring manager name.
Use case: A company hiring three SDRs is scaling outbound sales. They may need sales tools. A company hiring a Head of Data Engineering may need data infrastructure. Scrape the posting for company details, then find the relevant decision-maker's email through a data provider.
Review sites
G2, Capterra, TrustRadius and similar sites list companies that review software products, including competitor products.
What you find: Company names, reviewer names (sometimes), reviewer roles, company size.
Use case: Companies that reviewed a competitor's product are likely interested in your category. Build a list of those companies and target them.
Government and public records
Government databases contain business registrations, licensing data and regulatory filings.
Examples:
- SEC EDGAR (public company filings).
- State business registries (Secretary of State databases).
- FDA establishment registrations.
- Patent databases (USPTO, EPO).
- Government contract databases (FPDS, SAM.gov).
What you find: Company name, registered agent, address, filing details. Rarely includes email directly, but identifies companies to research further.
Social media and forums
LinkedIn, Twitter/X, Reddit and industry forums contain publicly posted information about professionals and companies.
Caution: LinkedIn's terms of service prohibit automated scraping. Other platforms have similar restrictions. Respect these terms. See Web Scraping Ethics and Best Practices.
The Scraping-to-Outreach Workflow
Step 1: Identify target sites
Based on your ICP, list the websites that contain information about your target companies or contacts.
Questions to answer:
- Where do your target companies list themselves?
- What directories or associations do they belong to?
- What conferences do they attend?
- What job boards do they post on?
Step 2: Check feasibility and legality
Before scraping any site:
- Read the Terms of Service.
- Check robots.txt.
- Assess whether the data is personal data under GDPR or other privacy laws.
- Determine your legal basis for collecting and using the data.
See Is Web Scraping Legal? and Scraping Ethics and Best Practices.
Step 3: Scrape the data
Use a scraping tool or script appropriate for the site's complexity.
For simple, static pages:
- Python with Beautiful Soup or Scrapy.
- No-code scrapers like Octoparse, ParseHub or WebScraper.io.
For JavaScript-rendered pages:
- Puppeteer or Playwright (headless browser).
- Selenium.
For paginated directories:
- Configure your scraper to follow pagination links.
- Respect rate limits (1-5 seconds between requests).
Step 4: Clean and structure the data
Scraped data is messy. Clean it before using it.
Common cleaning tasks:
- Remove HTML tags and formatting artefacts.
- Standardise company names (Inc., Inc, Incorporated are the same).
- Standardise job titles (VP, Vice President are the same).
- Remove incomplete records (no name, no company).
- Fix encoding issues (special characters, unicode).
Step 5: Extract email addresses
If the scraped data contains email addresses mixed with other text, extract them:
- Save the scraped data as a text file or CSV.
- Upload to Email Extractor.
- The tool extracts all email addresses and deduplicates them.
- Download the clean, unique email list.
Step 6: Find missing email addresses
If the scraped data includes names and companies but not email addresses (common with team pages and conference speakers), use an email finder to look them up.
Tools: Apollo.io, Hunter.io, Lusha, RocketReach, Snov.io.
Input: First name + Last name + Company domain.
Output: Verified email address.
Step 7: Enrich
Add firmographic and demographic data to scraped contacts:
- Company size, industry, revenue.
- Job title, seniority level.
- Technology stack.
- LinkedIn profile.
See Data Enrichment After Extraction.
Step 8: Verify
Verify every email address before sending. Scraped data has a higher-than-average bounce risk because:
- The data may be outdated (old team pages, past conference speakers).
- Email addresses may have been transcribed or formatted incorrectly.
- Some addresses may be obfuscated or partially visible on the source page.
See Best Email Verification Services.
Step 9: Segment and send
Segment the verified list by industry, role, company size or source. Write personalised messaging for each segment.
See How to Segment Your Email List and Cold Email Personalization at Scale.
Scraping vs Data Providers
| Factor | Web scraping | Data providers |
|---|---|---|
| Coverage | Whatever is publicly online | Pre-built databases, varies by provider |
| Niche data | Strong (specific directories, events) | Weaker for niche or small markets |
| Data freshness | As fresh as the source page | Updated periodically, can be stale |
| Setup effort | Moderate to high (coding or tool config) | Low (search and export) |
| Cost | Low (infrastructure costs) | Moderate to high (subscription fees) |
| Email availability | Often requires separate email lookup | Usually included |
| Compliance | You are responsible for GDPR, CAN-SPAM compliance | Provider handles some compliance |
| Scalability | Depends on infrastructure | Built-in |
Best approach: Use data providers as your primary source for broad prospecting. Use scraping to fill gaps: niche industries, specific events, companies not covered by providers, or data that changes frequently.
Common Mistakes
Scraping without a plan
Scraping everything you can find and hoping something useful emerges wastes time and resources. Start with your ICP, identify specific sites that list matching companies and contacts, and scrape with purpose.
Ignoring data quality
Scraped data requires more cleaning and verification than data from a provider. Budget time for data cleaning or your outreach will suffer from bad data.
Skipping verification
Scraped email addresses bounce at a higher rate than provider data. Always verify before sending.
Violating Terms of Service
Some sites explicitly prohibit scraping. Violating ToS risks legal action, IP blocking and reputational damage. Always check before scraping.
Not updating scraped lists
Data scraped six months ago is already decaying. Re-scrape or re-verify periodically. See Email List Decay Explained.