Scraping Online Directories for Leads: Techniques, Tools and Legal Considerations
On this page
Why Directories
Online directories aggregate business and professional contact information in structured, searchable formats. For lead generation, they represent concentrated sources of contact data organised by industry, location, profession or speciality.
Unlike scraping general websites, directories are purpose-built to display contact information. The data is structured, relatively consistent and often more current than scattered mentions across the web.
This guide covers how to extract lead data from online directories, the tools available and the legal and ethical boundaries.
Types of Online Directories
Business directories
General business directories:
- Yellow Pages / YP.com.
- Yelp (primarily local and service businesses).
- Better Business Bureau (BBB).
- Manta.
- Google Maps / Google Business Profiles.
Industry-specific directories:
- Thomas Net (manufacturing).
- Houzz (home services and design).
- Avvo (lawyers).
- Healthgrades / Zocdoc (healthcare providers).
- Clutch / G2 (B2B service providers).
- Capterra (software).
Professional directories
- Bar association directories (lawyers by jurisdiction).
- Medical board directories (physicians by state).
- CPA society directories (accountants by state).
- Engineering society directories.
- Real estate agent directories (Realtor.com, Zillow agent finder).
- Chamber of Commerce member directories.
Government and public directories
- SEC EDGAR (public company filings, officer information).
- State business entity searches (registered agents, officers).
- Government contractor databases (SAM.gov).
- Lobbying disclosure databases.
- Campaign finance databases.
Professional association directories
- Trade association member lists.
- Alumni directories.
- Conference attendee and speaker lists (when publicly posted).
- Award and recognition lists.
Scraping Techniques
Understanding directory structure
Before scraping any directory, study how it organises data:
Pagination. How does the directory split results across pages? URL parameters (?page=2), infinite scroll, "load more" buttons, or numbered page links.
Search and filtering. What search parameters are available? Location, category, keyword, alphabetical. These determine how you structure your scraping queries.
Detail pages. Does each listing have a detail page with more information than the list view? Often the list view shows name, category and city, while the detail page adds email, phone, website and description.
Rate limiting. How does the site respond to rapid requests? Blocking, CAPTCHAs, throttling, or no visible protection.
Manual extraction
For small directories (under 500 listings), manual extraction may be more efficient than building a scraper:
- Use the directory's search and filter to narrow results.
- Copy the results page content.
- Paste into a text file or spreadsheet.
- Upload to Email Extractor to extract email addresses from the pasted text.
When manual extraction makes sense:
- The directory is small (a few hundred listings).
- The data is needed once, not repeatedly.
- The directory has strong anti-scraping measures.
- You need the data quickly and do not have scraping tools set up.
Browser extensions
For medium-sized extractions (hundreds to low thousands of listings), browser extensions can automate the clicking and scrolling:
Web Scraper (Chrome extension). Create a sitemap defining the data you want to extract. The extension follows pagination and extracts data from each page.
Instant Data Scraper. AI-powered detection of tabular data on web pages. One-click extraction of visible data.
Data Miner. Pre-built recipes for common directories plus a custom recipe builder.
See Automating Web Data Collection for detailed setup instructions.
No-code scraping platforms
For larger or recurring extractions, no-code platforms handle pagination, anti-scraping measures and scheduling:
Browse AI. Train a robot by demonstrating what to extract. Handles pagination and dynamic content.
Octoparse. Visual workflow builder. Point-and-click to define extraction targets. Handles JavaScript-rendered pages.
ParseHub. Visual selector for defining data to extract. Handles interactive elements and pagination.
Apify. Pre-built actors for common directories. Custom actor development for specific needs.
Code-based scraping
For maximum control, flexibility and scale, write custom scrapers:
Python with BeautifulSoup (static pages):
import requests
from bs4 import BeautifulSoup
import time
def scrape_directory_page(url):
response = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'})
soup = BeautifulSoup(response.text, 'html.parser')
listings = []
for card in soup.select('.listing-card'):
listing = {
'name': card.select_one('.business-name').text.strip(),
'category': card.select_one('.category').text.strip(),
'location': card.select_one('.location').text.strip(),
}
email_link = card.select_one('a[href^="mailto:"]')
if email_link:
listing['email'] = email_link['href'].replace('mailto:', '')
website_link = card.select_one('a.website-link')
if website_link:
listing['website'] = website_link['href']
listings.append(listing)
return listings
base_url = 'https://example.com/directory?page={}'
all_listings = []
for page in range(1, 51):
listings = scrape_directory_page(base_url.format(page))
all_listings.extend(listings)
time.sleep(2)
Python with Playwright (JavaScript-rendered pages):
from playwright.sync_api import sync_playwright
import time
def scrape_dynamic_directory():
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto('https://example.com/directory')
while True:
page.wait_for_selector('.listing-card')
cards = page.query_selector_all('.listing-card')
for card in cards:
name = card.query_selector('.name').inner_text()
email_el = card.query_selector('a[href^="mailto:"]')
email = email_el.get_attribute('href').replace('mailto:', '') if email_el else None
print(f'{name}: {email}')
next_btn = page.query_selector('.next-page:not(.disabled)')
if not next_btn:
break
next_btn.click()
time.sleep(2)
browser.close()
Scrapy (large-scale crawling):
Scrapy is a full crawling framework for large-scale extractions (tens of thousands of pages). It handles concurrency, retries, data pipelines and export formats.
See Building a Scraping Pipeline for architecture guidance.
Data Cleaning After Extraction
Directory data is rarely clean enough to use directly. Common issues:
Formatting inconsistencies:
- Names in different formats (JOHN SMITH vs John Smith vs john smith).
- Phone numbers with different formatting (555-123-4567 vs (555) 123-4567 vs 5551234567).
- Addresses with inconsistent abbreviations (St. vs Street vs ST).
Incomplete records:
- Listings without email addresses.
- Listings with only a general info@ address.
- Listings missing phone numbers or websites.
Duplicate listings:
- Same business listed under multiple categories.
- Same business with different addresses (branches).
- Merged or acquired businesses with both old and new listings.
Outdated information:
- Closed businesses still listed.
- Old phone numbers and addresses.
- Former employees still listed as contacts.
Cleaning workflow
- Extract raw data from the directory.
- Standardise formatting: lowercase emails, title case names, consistent phone format.
- Extract and deduplicate emails: Upload the raw extraction output to Email Extractor to extract all email addresses and remove duplicates.
- Verify emails: Run through an email verification service to remove invalid addresses.
- Enrich incomplete records: Use a B2B data provider to fill in missing fields where possible.
- Remove obvious invalids: businesses with no contact information, clearly closed businesses, duplicate entries.
See Data Normalization Workflows and How to Clean an Email List.
Legal and Ethical Considerations
Legal landscape
The legality of web scraping varies by jurisdiction, the type of data, how the data is used and the terms of the website.
US legal framework:
- Computer Fraud and Abuse Act (CFAA). The hiQ Labs v. LinkedIn case (2022) established that scraping publicly accessible data is generally not a CFAA violation. However, circumventing technical barriers (login walls, CAPTCHAs) to access data may still violate the CFAA.
- Terms of Service. Violating a website's terms of service by scraping may constitute breach of contract but is generally not a criminal offence.
- State privacy laws (CCPA). Scraping personal data of California residents creates obligations under CCPA if you meet the threshold criteria.
EU legal framework:
- GDPR. Scraping personal data (names, email addresses) of EU residents requires a lawful basis for processing. Legitimate interest may apply for B2B prospecting, but you must conduct a legitimate interest assessment and provide an easy opt-out.
- Database Directive. The EU Database Directive protects databases as intellectual property. Extracting a "substantial part" of a database may infringe the database right.
Ethical guidelines
Legal permissibility does not mean ethical acceptability. Follow these guidelines:
Respect robots.txt. Check the site's robots.txt file for scraping restrictions. While not legally binding, ignoring it signals bad faith.
Rate limit your requests. Do not overwhelm the server. Add delays between requests (2-5 seconds minimum). Scrape during off-peak hours.
Do not circumvent access controls. If a directory requires login, subscription or payment to access data, do not bypass those controls.
Identify yourself. Use a descriptive User-Agent string that includes contact information.
Do not scrape personal data beyond business context. Business email, business phone and business address are appropriate for B2B prospecting. Personal email, home address and personal phone are not.
Honour opt-outs. When someone asks to be removed from your list, remove them immediately and permanently.
Comply with CAN-SPAM, GDPR and other applicable laws. Having the data does not give you permission to email anyone. Follow all applicable email laws when contacting scraped contacts.
See Is Web Scraping Legal and Scraping Ethics Best Practices.
Alternatives to Scraping
Official APIs
Many directories offer APIs for legitimate data access:
| Directory | API available | Access |
|---|---|---|
| Google Places | Yes | Paid, per-request pricing |
| Yelp | Yes | Free tier available |
| Crunchbase | Yes | Paid subscription |
| Yes (limited) | Partner programmes only | |
| BBB | No public API | N/A |
| Clutch | No public API | N/A |
APIs are always preferable to scraping when available. The data is structured, rate limits are defined and you are operating within the platform's terms.
Data providers
B2B data providers (ZoomInfo, Apollo, Cognism, Lusha) aggregate contact data from multiple sources including directories. Subscribing to a data provider may be more cost-effective than building and maintaining scrapers.
See B2B Data Provider Comparison.
Manual research
For high-value prospects, manual research yields the most accurate and current data:
- Visit the company's website for current team pages.
- Check LinkedIn for current titles and contact information.
- Call the company's main line and ask for the right contact.
- Attend industry events where prospects are present.
Partnership and referral
Some directories offer partnership or bulk data licensing agreements. If you need ongoing access to a directory's data, contact them about a commercial arrangement rather than scraping.
Building a Directory Scraping Workflow
Step 1: Identify target directories
List the directories that serve your target market. A company selling software to dentists might target:
- ADA (American Dental Association) member directory.
- State dental board licence databases.
- Healthgrades dental profiles.
- Google Maps dental practices by city.
- Yelp dental category.
Step 2: Evaluate each directory
For each directory, assess:
- How many listings are relevant?
- Is the data structured or unstructured?
- Are email addresses displayed or hidden?
- What anti-scraping measures are in place?
- What do the terms of service say?
- Is an API available?
Step 3: Choose the extraction method
| Directory size | Update frequency | Best method |
|---|---|---|
| Under 500 listings | One-time | Manual copy + Email Extractor |
| 500-5,000 listings | One-time | Browser extension |
| 500-5,000 listings | Recurring | No-code platform |
| 5,000+ listings | One-time or recurring | Custom scraper |
Step 4: Extract, clean and consolidate
- Run the extraction.
- Export results to CSV or text files.
- Upload to Email Extractor to extract and deduplicate email addresses across all directory extractions.
- Verify emails.
- Enrich with additional data if needed.
- Import into CRM or sales engagement platform.
Step 5: Maintain
Directories update regularly. Schedule re-extractions:
- Monthly for fast-changing directories (job boards, startup listings).
- Quarterly for moderate-change directories (professional directories, business listings).
- Annually for slow-change directories (government registrations, bar associations).