Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Scraping Online Directories for Leads: Techniques, Tools and Legal Considerations

On this page

Why Directories

Online directories aggregate business and professional contact information in structured, searchable formats. For lead generation, they represent concentrated sources of contact data organised by industry, location, profession or speciality.

Unlike scraping general websites, directories are purpose-built to display contact information. The data is structured, relatively consistent and often more current than scattered mentions across the web.

This guide covers how to extract lead data from online directories, the tools available and the legal and ethical boundaries.

Types of Online Directories

Business directories

General business directories:

  • Yellow Pages / YP.com.
  • Yelp (primarily local and service businesses).
  • Better Business Bureau (BBB).
  • Manta.
  • Google Maps / Google Business Profiles.

Industry-specific directories:

  • Thomas Net (manufacturing).
  • Houzz (home services and design).
  • Avvo (lawyers).
  • Healthgrades / Zocdoc (healthcare providers).
  • Clutch / G2 (B2B service providers).
  • Capterra (software).

Professional directories

  • Bar association directories (lawyers by jurisdiction).
  • Medical board directories (physicians by state).
  • CPA society directories (accountants by state).
  • Engineering society directories.
  • Real estate agent directories (Realtor.com, Zillow agent finder).
  • Chamber of Commerce member directories.

Government and public directories

  • SEC EDGAR (public company filings, officer information).
  • State business entity searches (registered agents, officers).
  • Government contractor databases (SAM.gov).
  • Lobbying disclosure databases.
  • Campaign finance databases.

Professional association directories

  • Trade association member lists.
  • Alumni directories.
  • Conference attendee and speaker lists (when publicly posted).
  • Award and recognition lists.

Scraping Techniques

Understanding directory structure

Before scraping any directory, study how it organises data:

Pagination. How does the directory split results across pages? URL parameters (?page=2), infinite scroll, "load more" buttons, or numbered page links.

Search and filtering. What search parameters are available? Location, category, keyword, alphabetical. These determine how you structure your scraping queries.

Detail pages. Does each listing have a detail page with more information than the list view? Often the list view shows name, category and city, while the detail page adds email, phone, website and description.

Rate limiting. How does the site respond to rapid requests? Blocking, CAPTCHAs, throttling, or no visible protection.

Manual extraction

For small directories (under 500 listings), manual extraction may be more efficient than building a scraper:

  1. Use the directory's search and filter to narrow results.
  2. Copy the results page content.
  3. Paste into a text file or spreadsheet.
  4. Upload to Email Extractor to extract email addresses from the pasted text.

When manual extraction makes sense:

  • The directory is small (a few hundred listings).
  • The data is needed once, not repeatedly.
  • The directory has strong anti-scraping measures.
  • You need the data quickly and do not have scraping tools set up.

Browser extensions

For medium-sized extractions (hundreds to low thousands of listings), browser extensions can automate the clicking and scrolling:

Web Scraper (Chrome extension). Create a sitemap defining the data you want to extract. The extension follows pagination and extracts data from each page.

Instant Data Scraper. AI-powered detection of tabular data on web pages. One-click extraction of visible data.

Data Miner. Pre-built recipes for common directories plus a custom recipe builder.

See Automating Web Data Collection for detailed setup instructions.

No-code scraping platforms

For larger or recurring extractions, no-code platforms handle pagination, anti-scraping measures and scheduling:

Browse AI. Train a robot by demonstrating what to extract. Handles pagination and dynamic content.

Octoparse. Visual workflow builder. Point-and-click to define extraction targets. Handles JavaScript-rendered pages.

ParseHub. Visual selector for defining data to extract. Handles interactive elements and pagination.

Apify. Pre-built actors for common directories. Custom actor development for specific needs.

Code-based scraping

For maximum control, flexibility and scale, write custom scrapers:

Python with BeautifulSoup (static pages):

import requests
from bs4 import BeautifulSoup
import time

def scrape_directory_page(url):
    response = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'})
    soup = BeautifulSoup(response.text, 'html.parser')

    listings = []
    for card in soup.select('.listing-card'):
        listing = {
            'name': card.select_one('.business-name').text.strip(),
            'category': card.select_one('.category').text.strip(),
            'location': card.select_one('.location').text.strip(),
        }

        email_link = card.select_one('a[href^="mailto:"]')
        if email_link:
            listing['email'] = email_link['href'].replace('mailto:', '')

        website_link = card.select_one('a.website-link')
        if website_link:
            listing['website'] = website_link['href']

        listings.append(listing)

    return listings

base_url = 'https://example.com/directory?page={}'
all_listings = []

for page in range(1, 51):
    listings = scrape_directory_page(base_url.format(page))
    all_listings.extend(listings)
    time.sleep(2)

Python with Playwright (JavaScript-rendered pages):

from playwright.sync_api import sync_playwright
import time

def scrape_dynamic_directory():
    with sync_playwright() as p:
        browser = p.chromium.launch()
        page = browser.new_page()

        page.goto('https://example.com/directory')

        while True:
            page.wait_for_selector('.listing-card')

            cards = page.query_selector_all('.listing-card')
            for card in cards:
                name = card.query_selector('.name').inner_text()
                email_el = card.query_selector('a[href^="mailto:"]')
                email = email_el.get_attribute('href').replace('mailto:', '') if email_el else None
                print(f'{name}: {email}')

            next_btn = page.query_selector('.next-page:not(.disabled)')
            if not next_btn:
                break
            next_btn.click()
            time.sleep(2)

        browser.close()

Scrapy (large-scale crawling):

Scrapy is a full crawling framework for large-scale extractions (tens of thousands of pages). It handles concurrency, retries, data pipelines and export formats.

See Building a Scraping Pipeline for architecture guidance.

Data Cleaning After Extraction

Directory data is rarely clean enough to use directly. Common issues:

Formatting inconsistencies:

  • Names in different formats (JOHN SMITH vs John Smith vs john smith).
  • Phone numbers with different formatting (555-123-4567 vs (555) 123-4567 vs 5551234567).
  • Addresses with inconsistent abbreviations (St. vs Street vs ST).

Incomplete records:

  • Listings without email addresses.
  • Listings with only a general info@ address.
  • Listings missing phone numbers or websites.

Duplicate listings:

  • Same business listed under multiple categories.
  • Same business with different addresses (branches).
  • Merged or acquired businesses with both old and new listings.

Outdated information:

  • Closed businesses still listed.
  • Old phone numbers and addresses.
  • Former employees still listed as contacts.

Cleaning workflow

  1. Extract raw data from the directory.
  2. Standardise formatting: lowercase emails, title case names, consistent phone format.
  3. Extract and deduplicate emails: Upload the raw extraction output to Email Extractor to extract all email addresses and remove duplicates.
  4. Verify emails: Run through an email verification service to remove invalid addresses.
  5. Enrich incomplete records: Use a B2B data provider to fill in missing fields where possible.
  6. Remove obvious invalids: businesses with no contact information, clearly closed businesses, duplicate entries.

See Data Normalization Workflows and How to Clean an Email List.

Legal landscape

The legality of web scraping varies by jurisdiction, the type of data, how the data is used and the terms of the website.

US legal framework:

  • Computer Fraud and Abuse Act (CFAA). The hiQ Labs v. LinkedIn case (2022) established that scraping publicly accessible data is generally not a CFAA violation. However, circumventing technical barriers (login walls, CAPTCHAs) to access data may still violate the CFAA.
  • Terms of Service. Violating a website's terms of service by scraping may constitute breach of contract but is generally not a criminal offence.
  • State privacy laws (CCPA). Scraping personal data of California residents creates obligations under CCPA if you meet the threshold criteria.

EU legal framework:

  • GDPR. Scraping personal data (names, email addresses) of EU residents requires a lawful basis for processing. Legitimate interest may apply for B2B prospecting, but you must conduct a legitimate interest assessment and provide an easy opt-out.
  • Database Directive. The EU Database Directive protects databases as intellectual property. Extracting a "substantial part" of a database may infringe the database right.

Ethical guidelines

Legal permissibility does not mean ethical acceptability. Follow these guidelines:

Respect robots.txt. Check the site's robots.txt file for scraping restrictions. While not legally binding, ignoring it signals bad faith.

Rate limit your requests. Do not overwhelm the server. Add delays between requests (2-5 seconds minimum). Scrape during off-peak hours.

Do not circumvent access controls. If a directory requires login, subscription or payment to access data, do not bypass those controls.

Identify yourself. Use a descriptive User-Agent string that includes contact information.

Do not scrape personal data beyond business context. Business email, business phone and business address are appropriate for B2B prospecting. Personal email, home address and personal phone are not.

Honour opt-outs. When someone asks to be removed from your list, remove them immediately and permanently.

Comply with CAN-SPAM, GDPR and other applicable laws. Having the data does not give you permission to email anyone. Follow all applicable email laws when contacting scraped contacts.

See Is Web Scraping Legal and Scraping Ethics Best Practices.

Alternatives to Scraping

Official APIs

Many directories offer APIs for legitimate data access:

Directory API available Access
Google Places Yes Paid, per-request pricing
Yelp Yes Free tier available
Crunchbase Yes Paid subscription
LinkedIn Yes (limited) Partner programmes only
BBB No public API N/A
Clutch No public API N/A

APIs are always preferable to scraping when available. The data is structured, rate limits are defined and you are operating within the platform's terms.

Data providers

B2B data providers (ZoomInfo, Apollo, Cognism, Lusha) aggregate contact data from multiple sources including directories. Subscribing to a data provider may be more cost-effective than building and maintaining scrapers.

See B2B Data Provider Comparison.

Manual research

For high-value prospects, manual research yields the most accurate and current data:

  • Visit the company's website for current team pages.
  • Check LinkedIn for current titles and contact information.
  • Call the company's main line and ask for the right contact.
  • Attend industry events where prospects are present.

Partnership and referral

Some directories offer partnership or bulk data licensing agreements. If you need ongoing access to a directory's data, contact them about a commercial arrangement rather than scraping.

Building a Directory Scraping Workflow

Step 1: Identify target directories

List the directories that serve your target market. A company selling software to dentists might target:

  • ADA (American Dental Association) member directory.
  • State dental board licence databases.
  • Healthgrades dental profiles.
  • Google Maps dental practices by city.
  • Yelp dental category.

Step 2: Evaluate each directory

For each directory, assess:

  • How many listings are relevant?
  • Is the data structured or unstructured?
  • Are email addresses displayed or hidden?
  • What anti-scraping measures are in place?
  • What do the terms of service say?
  • Is an API available?

Step 3: Choose the extraction method

Directory size Update frequency Best method
Under 500 listings One-time Manual copy + Email Extractor
500-5,000 listings One-time Browser extension
500-5,000 listings Recurring No-code platform
5,000+ listings One-time or recurring Custom scraper

Step 4: Extract, clean and consolidate

  1. Run the extraction.
  2. Export results to CSV or text files.
  3. Upload to Email Extractor to extract and deduplicate email addresses across all directory extractions.
  4. Verify emails.
  5. Enrich with additional data if needed.
  6. Import into CRM or sales engagement platform.

Step 5: Maintain

Directories update regularly. Schedule re-extractions:

  • Monthly for fast-changing directories (job boards, startup listings).
  • Quarterly for moderate-change directories (professional directories, business listings).
  • Annually for slow-change directories (government registrations, bar associations).

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)