Scraping Social Media for Business Contacts and Leads
On this page
Social Media as a Data Source
Social media platforms contain vast amounts of publicly shared business information: company profiles, employee directories, contact details, industry affiliations, technology usage and business relationships. This data is valuable for lead generation, market research and competitive analysis.
This guide covers what data is available on each platform, how to access it, the legal and ethical boundaries, and alternatives to scraping.
Platform-by-Platform Overview
Data available:
- Professional profiles (name, title, company, location, skills, experience).
- Company pages (industry, size, specialities, employee count, recent hires).
- Job postings (roles, requirements, technology stack).
- Content engagement (who comments on industry topics, who shares what).
- Group memberships (industry affiliations, interests).
- Connections and mutual connections.
Access methods:
| Method | Legality | Data quality | Scale |
|---|---|---|---|
| LinkedIn Sales Navigator | Permitted (official tool) | High | Limited by subscription |
| LinkedIn API (official) | Permitted (for approved apps) | High | Rate-limited |
| Manual research | Permitted | High | Very limited |
| Third-party data providers (Apollo, ZoomInfo) | Grey area (they collect, you buy) | Medium-high | Large |
| Direct scraping | Prohibited by ToS; court rulings mixed | High | Unlimited technically |
| Chrome extension tools (Dux-Soup, Expandi) | Prohibited by ToS; risk of account ban | High | Limited by daily caps |
LinkedIn's position: LinkedIn actively prohibits scraping and has pursued legal action against scrapers. The landmark hiQ v. LinkedIn case (2022) established that scraping publicly available data may not violate the CFAA, but LinkedIn has continued to enforce its ToS through account restrictions and legal threats.
Recommendation: Use LinkedIn Sales Navigator for research and prospecting. Use official data providers for contact data. Do not scrape LinkedIn directly.
Twitter/X
Data available:
- Public profiles (name, bio, location, links, follower count).
- Tweets and replies (content, engagement, timing).
- Followers and following lists.
- Lists (curated groups of users by topic).
- Spaces participation.
Access methods:
| Method | Legality | Data quality | Scale |
|---|---|---|---|
| X API (official) | Permitted (paid tiers) | High | Rate-limited by tier |
| Twitter/X search | Permitted | Medium | Manual |
| Third-party tools (Followerwonk, SparkToro) | Permitted (use API) | Medium-high | Limited by API |
| Direct scraping | Prohibited by ToS | Medium | Variable |
Business use cases:
- Find industry thought leaders and influencers.
- Monitor competitor activity and positioning.
- Identify prospects discussing relevant topics.
- Discover company employees and their interests.
- Research industry trends and conversations.
Facebook and Instagram
Data available:
- Business pages (company info, contact details, reviews).
- Groups (member lists on public groups, discussions).
- Ad library (competitor advertising).
- Instagram business profiles (contact info, category).
Access methods:
| Method | Legality | Scale |
|---|---|---|
| Meta Business Suite / Graph API | Permitted (for page owners and approved apps) | Limited |
| Ad Library API | Permitted | Moderate |
| Manual research | Permitted | Very limited |
| Direct scraping | Prohibited by ToS; violates terms aggressively enforced | Variable |
Business use cases:
- Competitor ad monitoring (Ad Library).
- Local business contact information.
- Customer sentiment from reviews and comments.
- Industry group discussions and member identification.
GitHub
Data available:
- Developer profiles (name, email, location, company, repositories).
- Organisation profiles (members, repositories, technology stack).
- Repository data (languages, dependencies, contributors).
- Issue and pull request discussions.
Access methods:
| Method | Legality | Scale |
|---|---|---|
| GitHub API (official) | Permitted | Rate-limited (5,000/hour authenticated) |
| GitHub search | Permitted | Moderate |
| Git log (commit history) | Public data | Per repository |
Business use cases:
- Developer recruiting (find developers by language, contribution level, location).
- Technology adoption tracking (which companies use which frameworks).
- Competitor engineering analysis (open-source activity, hiring, technology choices).
- Finding developer emails from commit history.
Note on emails from commits: Git commits often contain the committer's email address. These are publicly visible in open-source repositories. While technically public data, using scraped commit emails for cold outreach is generally unwelcome in the developer community.
Data available:
- Subreddit discussions (topics, opinions, recommendations).
- User profiles (post history, subreddit participation).
- Awards and engagement data.
Access methods:
| Method | Legality | Scale |
|---|---|---|
| Reddit API (official) | Permitted (with restrictions since 2023 API changes) | Rate-limited |
| Reddit search | Permitted | Moderate |
| Manual research | Permitted | Limited |
Business use cases:
- Market research (what do people say about your category?).
- Competitive intelligence (competitor complaints and praise).
- Product feedback (subreddit mentions of your product or competitors).
- Community identification (which subreddits discuss your industry?).
Reddit is generally poor for direct contact extraction but excellent for market research and competitive intelligence.
Data Extraction Techniques
API-first approach (recommended)
APIs are the preferred method for accessing social media data because:
- Access is explicitly permitted.
- Data is structured and clean.
- Rate limits are documented.
- No risk of account bans.
- Compliance with platform policies.
API workflow:
- Register for API access on the platform.
- Obtain API credentials (API key, OAuth tokens).
- Build queries to extract the data you need.
- Respect rate limits (implement backoff and retry logic).
- Store data responsibly (comply with platform data use policies).
Browser extension approach
For platforms where API access is limited (especially LinkedIn), browser extensions can automate manual research:
| Tool | Platform | What it does |
|---|---|---|
| LinkedIn Sales Navigator | Official prospecting tool with search and save | |
| Apollo Chrome Extension | LinkedIn, websites | Find contact data while browsing profiles |
| Lusha | LinkedIn, websites | Reveal contact data on profiles |
| Hunter Chrome Extension | Websites | Find email addresses on company websites |
| Clearbit Connect (discontinued) | Historical Gmail tool | Sunset on 30 April 2025; evaluate current alternatives separately |
These extensions automate what a user could do manually: view a profile and extract visible data. They operate in a grey area: faster than manual but slower than direct scraping.
Export approach
Many platforms allow you to export your own data:
| Platform | What you can export |
|---|---|
| Your connections (name, email, company, title) | |
| Twitter/X | Your followers, following, tweets |
| Your friends, page followers, group members (if admin) | |
| Your followers and following | |
| Mailchimp, HubSpot, etc. | Subscriber lists |
Exporting your own LinkedIn connections is the most legitimate way to get contact data from LinkedIn. Go to Settings > Data Privacy > Get a copy of your data > Connections. The export includes name, email (if shared), company, title and connection date.
After exporting, upload the CSV to Email Extractor to extract and deduplicate the email addresses from the export files.
Data Enrichment
Social media data is often incomplete. Enrichment fills the gaps:
Workflow:
- Extract or export base data from social media (name, company, title).
- Use an email finder tool to discover email addresses. See Email Finder Tool Comparison.
- Use a data provider for phone numbers, company data, technology data.
- Consolidate and deduplicate. Upload all sources to Email Extractor.
- Verify email addresses before outreach.
Enrichment providers:
| Provider | What they add |
|---|---|
| Apollo.io | Email, phone, company data, technology |
| ZoomInfo | Email, phone, company data, intent data |
| Clearbit (HubSpot) | Company data, technology, web traffic |
| Clay | Aggregated enrichment from 50+ sources |
| FullContact | Identity resolution, social profiles |
Legal and Ethical Considerations
Terms of service
Every major social media platform prohibits scraping in its terms of service:
| Platform | ToS prohibition | Enforcement |
|---|---|---|
| Explicitly prohibits scraping and automated access | Aggressive (account bans, legal action) | |
| Twitter/X | Prohibits scraping; aggressive API pricing since 2023 | Account suspension, legal threats |
| Facebook/Instagram | Prohibits scraping; CFAA claims against scrapers | Account bans, legal action |
| GitHub | Allows API access; prohibits excessive scraping | Rate limiting, account suspension |
| Restricts API access since 2023 | API pricing, rate limiting |
Legal landscape
- hiQ v. LinkedIn (2022): The Ninth Circuit ruled that scraping publicly available data may not violate the CFAA. However, this does not override ToS or other legal claims.
- GDPR (EU): Personal data scraped from social media is still personal data. Processing requires a legal basis (legitimate interest is the most commonly claimed basis for B2B).
- CCPA (California): Publicly available information is generally excluded from CCPA's definition of personal information, but the exclusion is narrow.
- Platform-specific data policies: Platforms may require that data obtained through their APIs is not used for certain purposes (e.g., advertising, reselling).
Ethical guidelines
Even where scraping may be technically legal:
- Respect opt-out signals (private accounts, do-not-contact requests).
- Do not contact individuals about sensitive information revealed on social media.
- Do not misrepresent how you obtained contact information.
- Rate limit requests to avoid impacting platform performance.
- Do not redistribute or resell scraped data.
- Consider whether the individual would expect to be contacted based on what they shared.
See Is Web Scraping Legal and Scraping Ethics Best Practices.
Alternatives to Scraping
| Alternative | What it provides | Compliance |
|---|---|---|
| Official APIs | Structured data within platform rules | Fully compliant |
| B2B data providers (Apollo, ZoomInfo) | Contact data already collected and verified | Generally compliant |
| LinkedIn Sales Navigator | Advanced search and prospecting | Fully compliant |
| Intent data providers (Bombora, G2) | Buying signals without scraping | Fully compliant |
| Industry directories and databases | Business listings and contact data | Generally compliant |
| Your own data exports | Your connections and subscribers | Fully compliant |
| Email finder tools | Email addresses from name + company | Varies by provider |
| Event attendee lists | Contacts who opted in at events | Compliant with consent |