Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Web Scraping for Market Research: Competitive Intelligence, Pricing and Trend Analysis

On this page

Scraping Beyond Lead Generation

Most guides on web scraping focus on extracting contact information for sales prospecting. But scraping has broader applications for market research: monitoring competitors, tracking pricing, analysing reviews, understanding market trends and gathering data for strategic decisions.

This guide covers how to use web scraping for market research purposes.

Competitive Intelligence

Monitoring competitor websites

What to track:

  • Product pages (new features, removed features, product changes).
  • Pricing pages (price changes, new tiers, new packaging).
  • Blog and content (topics, frequency, positioning).
  • Careers pages (what roles they are hiring for signals strategic direction).
  • Case studies (which industries and use cases they highlight).
  • Press releases and news (partnerships, funding, executive changes).

Implementation:

Set up automated monitoring that checks competitor pages on a schedule and alerts you to changes.

Simple approach (no coding):

  • Use a change detection service (Visualping, ChangeTower, Distill.io).
  • Monitor specific URLs (pricing page, product page, careers page).
  • Receive email alerts when content changes.

Code-based approach:

import requests
import hashlib
import json
from datetime import datetime

def check_page(url, previous_hash):
    response = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'})
    current_hash = hashlib.md5(response.text.encode()).hexdigest()

    if current_hash != previous_hash:
        return {
            'url': url,
            'changed': True,
            'timestamp': datetime.now().isoformat(),
            'new_hash': current_hash
        }
    return {'url': url, 'changed': False, 'new_hash': current_hash}

competitors = {
    'Competitor A Pricing': 'https://example.com/pricing',
    'Competitor A Careers': 'https://example.com/careers',
    'Competitor B Pricing': 'https://example.org/pricing',
}

# Load previous hashes from storage
with open('hashes.json', 'r') as f:
    previous_hashes = json.load(f)

for name, url in competitors.items():
    result = check_page(url, previous_hashes.get(name, ''))
    if result['changed']:
        print(f'CHANGE DETECTED: {name}')
        # Send alert email or Slack notification
    previous_hashes[name] = result['new_hash']

# Save updated hashes
with open('hashes.json', 'w') as f:
    json.dump(previous_hashes, f)

Competitor content analysis

Scraping competitor blogs and content hubs reveals their content strategy:

Data to extract:

  • Article titles and URLs.
  • Publication dates.
  • Categories and tags.
  • Author names.
  • Word count (approximate content depth).
  • Social sharing counts (when visible).

Analysis questions:

  • What topics do they publish about most frequently?
  • What topics have they started covering recently?
  • How often do they publish?
  • What content format do they use (listicles, how-to guides, case studies, thought leadership)?
  • Which topics get the most engagement?

Use for your own strategy:

  • Identify content gaps (topics they cover that you do not).
  • Identify over-served topics (where adding another article has diminishing returns).
  • Understand their positioning (how they describe the problem and their solution).
  • Track shifts in messaging (new language, new value propositions).

Pricing Intelligence

Why pricing data matters

Understanding competitor pricing helps you:

  • Position your own pricing competitively.
  • Identify market opportunities (underserved price points).
  • Track price trends over time.
  • Respond to competitor price changes.
  • Support pricing decisions with market data.

What to scrape

SaaS and subscription businesses:

  • Plan names and descriptions.
  • Pricing for each plan (monthly and annual).
  • Features included in each plan.
  • Add-on pricing.
  • Enterprise "contact us" indicators (which signal custom pricing).
  • Free tier or trial availability.

E-commerce and retail:

  • Product prices (regular and sale).
  • Shipping costs.
  • Bundle and quantity discounts.
  • Price history (track changes over time).
  • Stock availability.

Services and professional fees:

  • Published rate cards.
  • Package pricing.
  • Hourly or project rate ranges.

Building a pricing tracker

Components:

  1. Scraper. Extracts pricing data from competitor websites on a schedule.
  2. Storage. Database to store pricing snapshots over time.
  3. Analysis. Comparison of current prices, historical trends, and alerts for changes.
  4. Reporting. Dashboard or report for the pricing and product team.

Example: SaaS pricing tracker

import requests
from bs4 import BeautifulSoup
import sqlite3
from datetime import date

def scrape_pricing(url, selectors):
    response = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'})
    soup = BeautifulSoup(response.text, 'html.parser')

    plans = []
    for plan_card in soup.select(selectors['plan_card']):
        plan = {
            'name': plan_card.select_one(selectors['plan_name']).text.strip(),
            'price': plan_card.select_one(selectors['price']).text.strip(),
            'features': [f.text.strip() for f in plan_card.select(selectors['features'])],
            'date': date.today().isoformat()
        }
        plans.append(plan)
    return plans

# Store in database for historical tracking
conn = sqlite3.connect('pricing.db')
cursor = conn.cursor()
cursor.execute('''
    CREATE TABLE IF NOT EXISTS pricing (
        id INTEGER PRIMARY KEY,
        competitor TEXT,
        plan_name TEXT,
        price TEXT,
        features TEXT,
        scrape_date TEXT
    )
''')

# Scrape and store
# Each competitor needs its own CSS selectors
competitors = {
    'Competitor A': {
        'url': 'https://example.com/pricing',
        'selectors': {
            'plan_card': '.pricing-card',
            'plan_name': '.plan-title',
            'price': '.price-amount',
            'features': '.feature-item'
        }
    }
}

for name, config in competitors.items():
    plans = scrape_pricing(config['url'], config['selectors'])
    for plan in plans:
        cursor.execute(
            'INSERT INTO pricing (competitor, plan_name, price, features, scrape_date) VALUES (?, ?, ?, ?, ?)',
            (name, plan['name'], plan['price'], str(plan['features']), plan['date'])
        )

conn.commit()
conn.close()

Pricing data challenges

  • Dynamic pricing. Some sites show different prices based on location, device or user history. Use consistent scraping conditions (same IP region, clean cookies).
  • JavaScript-rendered pricing. Many SaaS pricing pages render prices with JavaScript. Use Playwright or Selenium instead of BeautifulSoup.
  • Pricing toggles. Monthly vs annual pricing, different currencies, different billing cycles. Capture all variants.
  • Gated pricing. Enterprise pricing behind a "contact us" form. You cannot scrape this; note the existence of custom pricing.

Review and Sentiment Analysis

Scraping reviews

Customer reviews on third-party platforms provide unfiltered market intelligence.

Sources:

  • G2, Capterra, TrustRadius (B2B software).
  • Yelp, Google Reviews (local businesses).
  • Amazon, product review sites (consumer products).
  • Glassdoor (employer reviews, useful for competitive intelligence on culture and strategy).
  • App Store, Google Play (mobile apps).

Data to extract:

  • Review text.
  • Rating (stars).
  • Review date.
  • Reviewer role or company (when available).
  • Product/feature mentioned.
  • Pros and cons sections.

Analysis approaches

Quantitative analysis:

  • Average rating over time (is competitor satisfaction trending up or down?).
  • Rating distribution (what percentage are 1-star vs 5-star?).
  • Volume of reviews over time (is the product growing?).
  • Review velocity (how many reviews per month?).

Qualitative analysis:

  • Common complaints (what do customers dislike?).
  • Common praises (what do customers love?).
  • Feature requests (what do customers wish the product did?).
  • Competitive mentions (which competitors are mentioned in reviews?).
  • Migration patterns ("switched from [Competitor]" or "moving to [Competitor]").

Use for your business:

  • Identify competitor weaknesses to exploit in positioning.
  • Understand what customers value most.
  • Discover unmet needs in the market.
  • Validate or challenge your own feature priorities.
  • Create marketing content addressing competitor complaints.

Job Posting Analysis

What job postings reveal

Competitor job postings are some of the most revealing public signals of strategic direction.

What to extract:

  • Job title and department.
  • Location (or remote).
  • Required skills and technologies.
  • Team size mentions.
  • Product or project descriptions.
  • Seniority level.
  • Posting date.

What it reveals:

Signal Interpretation
Hiring many engineers Investing in product development
Hiring data scientists Building analytics or AI capabilities
Hiring enterprise sales reps Moving upmarket
Hiring in a new city/country Expanding geographically
Listing a new technology in requirements Adopting a new tech stack
Hiring for a product area they don't have yet Building a new product or feature
Hiring a VP of a function they didn't have Formalising a new business area

Implementation

Sources to scrape:

  • Company career pages.
  • LinkedIn job postings.
  • Indeed, Glassdoor, ZipRecruiter.
  • Specialised job boards (AngelList/Wellfound for startups, Dice for tech).

Tracking approach:

  1. Scrape competitor career pages weekly.
  2. Track new postings, filled postings and removed postings.
  3. Categorise by function (engineering, sales, marketing, operations).
  4. Track trends (increasing hiring in a department signals investment).
  5. Note specific technologies and skills mentioned.

Trend Analysis

Industry trend monitoring

Scraping industry-relevant data sources reveals market trends:

Sources:

  • Industry news sites and blogs.
  • Conference agendas and speaker lists (what topics are trending).
  • Academic and research publications.
  • Patent filings.
  • Regulatory filings.
  • Social media hashtags and discussions.

Technology trend monitoring

Sources:

  • Stack Overflow trends (technology adoption).
  • GitHub trending repositories and stars.
  • npm/PyPI download statistics (library popularity).
  • Hacker News and Reddit discussions.
  • Technology survey results (Stack Overflow Developer Survey, etc.).

Market sizing data

Sources:

  • Industry reports (summary data often available on report landing pages).
  • Government statistics (census, trade data, industry statistics).
  • Public company filings (revenue, growth rates, segment data).
  • Press releases with market data.
  • Association and trade group published statistics.

Using Scraped Data with Email Outreach

Market research data can inform email outreach:

Competitive intelligence in outreach:

  • Reference a prospect's current tool and its known limitations.
  • Address specific pain points identified in competitor reviews.
  • Reference industry trends relevant to the prospect's business.

Pricing intelligence in outreach:

  • Position your pricing against market benchmarks.
  • Highlight value differences at similar price points.

Contact data from research:

During market research, you may collect contact information incidentally (from press releases, conference speaker lists, published case studies). Consolidate this with your existing contact data:

  1. Export contacts gathered from research sources.
  2. Upload to Email Extractor to extract and deduplicate email addresses.
  3. Verify addresses before adding to outreach lists.
  4. Ensure you have a lawful basis for contacting them (legitimate interest for B2B in most jurisdictions; see compliance guides).

Ethical scraping for research

  • Rate limit requests. Do not overload target servers. 2-5 second delays between requests.
  • Respect robots.txt. While not legally binding, it is the site's stated preference.
  • Do not bypass access controls. If data is behind a paywall or login, do not circumvent it.
  • Use data for analysis, not republishing. Scraping for competitive intelligence (internal analysis) is different from scraping to republish content.
  • Attribute sources. When publishing research based on scraped data, cite the sources.

Legal framework

  • Public data. Scraping publicly available data for research purposes is generally permissible under US law (per hiQ v. LinkedIn).
  • Terms of service. Some sites prohibit scraping in their ToS. Violating ToS may create contractual liability.
  • GDPR and privacy. If your research involves personal data of EU residents, GDPR applies. Market research on pricing, products and trends typically does not involve personal data.
  • Copyright. Scraping copyrighted content (articles, images) for republishing may violate copyright law. Scraping for analysis (facts, data, trends) is generally permissible.

See Is Web Scraping Legal and Scraping Ethics Best Practices.

Tools

Tool Best for Complexity
Visualping / ChangeTower Simple page change monitoring No code
Browse AI / Octoparse Structured data extraction, recurring scraping No code
Beautiful Soup (Python) Static pages, custom extraction Code required
Playwright / Selenium JavaScript-rendered pages Code required
Scrapy Large-scale, multi-page crawling Code required
Apify Pre-built scrapers for common sites Low code

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)