Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Scraping Social Media for Business Contacts and Leads

On this page

Social Media as a Data Source

Social media platforms contain vast amounts of publicly shared business information: company profiles, employee directories, contact details, industry affiliations, technology usage and business relationships. This data is valuable for lead generation, market research and competitive analysis.

This guide covers what data is available on each platform, how to access it, the legal and ethical boundaries, and alternatives to scraping.

Platform-by-Platform Overview

LinkedIn

Data available:

  • Professional profiles (name, title, company, location, skills, experience).
  • Company pages (industry, size, specialities, employee count, recent hires).
  • Job postings (roles, requirements, technology stack).
  • Content engagement (who comments on industry topics, who shares what).
  • Group memberships (industry affiliations, interests).
  • Connections and mutual connections.

Access methods:

Method Legality Data quality Scale
LinkedIn Sales Navigator Permitted (official tool) High Limited by subscription
LinkedIn API (official) Permitted (for approved apps) High Rate-limited
Manual research Permitted High Very limited
Third-party data providers (Apollo, ZoomInfo) Grey area (they collect, you buy) Medium-high Large
Direct scraping Prohibited by ToS; court rulings mixed High Unlimited technically
Chrome extension tools (Dux-Soup, Expandi) Prohibited by ToS; risk of account ban High Limited by daily caps

LinkedIn's position: LinkedIn actively prohibits scraping and has pursued legal action against scrapers. The landmark hiQ v. LinkedIn case (2022) established that scraping publicly available data may not violate the CFAA, but LinkedIn has continued to enforce its ToS through account restrictions and legal threats.

Recommendation: Use LinkedIn Sales Navigator for research and prospecting. Use official data providers for contact data. Do not scrape LinkedIn directly.

Twitter/X

Data available:

  • Public profiles (name, bio, location, links, follower count).
  • Tweets and replies (content, engagement, timing).
  • Followers and following lists.
  • Lists (curated groups of users by topic).
  • Spaces participation.

Access methods:

Method Legality Data quality Scale
X API (official) Permitted (paid tiers) High Rate-limited by tier
Twitter/X search Permitted Medium Manual
Third-party tools (Followerwonk, SparkToro) Permitted (use API) Medium-high Limited by API
Direct scraping Prohibited by ToS Medium Variable

Business use cases:

  • Find industry thought leaders and influencers.
  • Monitor competitor activity and positioning.
  • Identify prospects discussing relevant topics.
  • Discover company employees and their interests.
  • Research industry trends and conversations.

Facebook and Instagram

Data available:

  • Business pages (company info, contact details, reviews).
  • Groups (member lists on public groups, discussions).
  • Ad library (competitor advertising).
  • Instagram business profiles (contact info, category).

Access methods:

Method Legality Scale
Meta Business Suite / Graph API Permitted (for page owners and approved apps) Limited
Ad Library API Permitted Moderate
Manual research Permitted Very limited
Direct scraping Prohibited by ToS; violates terms aggressively enforced Variable

Business use cases:

  • Competitor ad monitoring (Ad Library).
  • Local business contact information.
  • Customer sentiment from reviews and comments.
  • Industry group discussions and member identification.

GitHub

Data available:

  • Developer profiles (name, email, location, company, repositories).
  • Organisation profiles (members, repositories, technology stack).
  • Repository data (languages, dependencies, contributors).
  • Issue and pull request discussions.

Access methods:

Method Legality Scale
GitHub API (official) Permitted Rate-limited (5,000/hour authenticated)
GitHub search Permitted Moderate
Git log (commit history) Public data Per repository

Business use cases:

  • Developer recruiting (find developers by language, contribution level, location).
  • Technology adoption tracking (which companies use which frameworks).
  • Competitor engineering analysis (open-source activity, hiring, technology choices).
  • Finding developer emails from commit history.

Note on emails from commits: Git commits often contain the committer's email address. These are publicly visible in open-source repositories. While technically public data, using scraped commit emails for cold outreach is generally unwelcome in the developer community.

Reddit

Data available:

  • Subreddit discussions (topics, opinions, recommendations).
  • User profiles (post history, subreddit participation).
  • Awards and engagement data.

Access methods:

Method Legality Scale
Reddit API (official) Permitted (with restrictions since 2023 API changes) Rate-limited
Reddit search Permitted Moderate
Manual research Permitted Limited

Business use cases:

  • Market research (what do people say about your category?).
  • Competitive intelligence (competitor complaints and praise).
  • Product feedback (subreddit mentions of your product or competitors).
  • Community identification (which subreddits discuss your industry?).

Reddit is generally poor for direct contact extraction but excellent for market research and competitive intelligence.

Data Extraction Techniques

API-first approach (recommended)

APIs are the preferred method for accessing social media data because:

  • Access is explicitly permitted.
  • Data is structured and clean.
  • Rate limits are documented.
  • No risk of account bans.
  • Compliance with platform policies.

API workflow:

  1. Register for API access on the platform.
  2. Obtain API credentials (API key, OAuth tokens).
  3. Build queries to extract the data you need.
  4. Respect rate limits (implement backoff and retry logic).
  5. Store data responsibly (comply with platform data use policies).

Browser extension approach

For platforms where API access is limited (especially LinkedIn), browser extensions can automate manual research:

Tool Platform What it does
LinkedIn Sales Navigator LinkedIn Official prospecting tool with search and save
Apollo Chrome Extension LinkedIn, websites Find contact data while browsing profiles
Lusha LinkedIn, websites Reveal contact data on profiles
Hunter Chrome Extension Websites Find email addresses on company websites
Clearbit Connect (discontinued) Historical Gmail tool Sunset on 30 April 2025; evaluate current alternatives separately

These extensions automate what a user could do manually: view a profile and extract visible data. They operate in a grey area: faster than manual but slower than direct scraping.

Export approach

Many platforms allow you to export your own data:

Platform What you can export
LinkedIn Your connections (name, email, company, title)
Twitter/X Your followers, following, tweets
Facebook Your friends, page followers, group members (if admin)
Instagram Your followers and following
Mailchimp, HubSpot, etc. Subscriber lists

Exporting your own LinkedIn connections is the most legitimate way to get contact data from LinkedIn. Go to Settings > Data Privacy > Get a copy of your data > Connections. The export includes name, email (if shared), company, title and connection date.

After exporting, upload the CSV to Email Extractor to extract and deduplicate the email addresses from the export files.

Data Enrichment

Social media data is often incomplete. Enrichment fills the gaps:

Workflow:

  1. Extract or export base data from social media (name, company, title).
  2. Use an email finder tool to discover email addresses. See Email Finder Tool Comparison.
  3. Use a data provider for phone numbers, company data, technology data.
  4. Consolidate and deduplicate. Upload all sources to Email Extractor.
  5. Verify email addresses before outreach.

Enrichment providers:

Provider What they add
Apollo.io Email, phone, company data, technology
ZoomInfo Email, phone, company data, intent data
Clearbit (HubSpot) Company data, technology, web traffic
Clay Aggregated enrichment from 50+ sources
FullContact Identity resolution, social profiles

Terms of service

Every major social media platform prohibits scraping in its terms of service:

Platform ToS prohibition Enforcement
LinkedIn Explicitly prohibits scraping and automated access Aggressive (account bans, legal action)
Twitter/X Prohibits scraping; aggressive API pricing since 2023 Account suspension, legal threats
Facebook/Instagram Prohibits scraping; CFAA claims against scrapers Account bans, legal action
GitHub Allows API access; prohibits excessive scraping Rate limiting, account suspension
Reddit Restricts API access since 2023 API pricing, rate limiting

Legal landscape

  • hiQ v. LinkedIn (2022): The Ninth Circuit ruled that scraping publicly available data may not violate the CFAA. However, this does not override ToS or other legal claims.
  • GDPR (EU): Personal data scraped from social media is still personal data. Processing requires a legal basis (legitimate interest is the most commonly claimed basis for B2B).
  • CCPA (California): Publicly available information is generally excluded from CCPA's definition of personal information, but the exclusion is narrow.
  • Platform-specific data policies: Platforms may require that data obtained through their APIs is not used for certain purposes (e.g., advertising, reselling).

Ethical guidelines

Even where scraping may be technically legal:

  • Respect opt-out signals (private accounts, do-not-contact requests).
  • Do not contact individuals about sensitive information revealed on social media.
  • Do not misrepresent how you obtained contact information.
  • Rate limit requests to avoid impacting platform performance.
  • Do not redistribute or resell scraped data.
  • Consider whether the individual would expect to be contacted based on what they shared.

See Is Web Scraping Legal and Scraping Ethics Best Practices.

Alternatives to Scraping

Alternative What it provides Compliance
Official APIs Structured data within platform rules Fully compliant
B2B data providers (Apollo, ZoomInfo) Contact data already collected and verified Generally compliant
LinkedIn Sales Navigator Advanced search and prospecting Fully compliant
Intent data providers (Bombora, G2) Buying signals without scraping Fully compliant
Industry directories and databases Business listings and contact data Generally compliant
Your own data exports Your connections and subscribers Fully compliant
Email finder tools Email addresses from name + company Varies by provider
Event attendee lists Contacts who opted in at events Compliant with consent

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)