Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Scraping E-Commerce Product Data for Business Intelligence

On this page

Business Uses for E-Commerce Data

E-commerce product data powers a wide range of business intelligence applications:

Use case What data you need Business value
Competitive pricing Competitor prices for the same or similar products Optimise your own pricing strategy
Assortment analysis Competitor product catalogues Identify gaps in your offering
MAP (Minimum Advertised Price) monitoring Reseller prices for your products Enforce pricing agreements
Market trend analysis Category-level product and pricing data over time Identify growing or declining categories
Product research Reviews, ratings, Q&A data Understand customer needs and pain points
Lead generation Seller/merchant contact information Reach potential customers or partners
Supply chain intelligence Product availability, shipping times, stock status Monitor supply chain health
Demand forecasting Review velocity, best-seller rankings, stock changes Predict demand trends
Brand monitoring Your product listings across retailers Ensure accuracy; detect counterfeits
New product discovery Newly listed products in your category Stay ahead of competitor launches

Data Sources and Access Methods

Major e-commerce platforms

Platform API available Scraping policy Best data access method
Amazon Product Advertising API (PA API) Prohibited in terms of service PA API (affiliate programme); Amazon SP-API (sellers)
Shopify stores Storefront API (if enabled by merchant) Varies by store Structured data (JSON-LD); Storefront API where available
eBay Browse API, Finding API Restricted in terms of service eBay APIs (developer programme)
Walmart Affiliate API Prohibited in terms of service Walmart API (affiliate programme)
Etsy Open API Restricted Etsy API (developer programme)
Best Buy Products API Restricted Best Buy API (developer programme)
Target No public API Prohibited No official access
AliExpress Affiliate API Restricted AliExpress API (affiliate)

Structured data on product pages

Many e-commerce sites include structured data (JSON-LD, Microdata, RDFa) in their HTML that is designed for search engines. This data is machine-readable and often includes:

Schema.org type Fields typically included Where to find
Product Name, description, SKU, brand, image, category JSON-LD in page <head> or <body>
Offer Price, currency, availability, seller, condition Nested within Product
AggregateRating Rating value, review count Nested within Product
Review Author, rating, review body, date Nested within Product or separate
BreadcrumbList Category hierarchy JSON-LD in page <head>

Extracting structured data with Python

import requests
from bs4 import BeautifulSoup
import json

def extract_structured_data(url, headers=None):
    """Extract JSON-LD structured data from a product page."""
    if headers is None:
        headers = {
            'User-Agent': (
                'Mozilla/5.0 (compatible; '
                'DataResearchBot/1.0; '
                '+https://example.com/bot)'
            )
        }

    response = requests.get(url, headers=headers, timeout=30)
    response.raise_for_status()

    soup = BeautifulSoup(response.text, 'html.parser')
    structured_data = []

    # Find all JSON-LD blocks
    for script in soup.find_all(
        'script', type='application/ld+json'
    ):
        try:
            data = json.loads(script.string)
            structured_data.append(data)
        except (json.JSONDecodeError, TypeError):
            continue

    # Filter for product data
    products = []
    for data in structured_data:
        if isinstance(data, list):
            for item in data:
                if item.get('@type') == 'Product':
                    products.append(item)
        elif isinstance(data, dict):
            if data.get('@type') == 'Product':
                products.append(data)
            # Check @graph
            if '@graph' in data:
                for item in data['@graph']:
                    if item.get('@type') == 'Product':
                        products.append(item)

    return products


def parse_product(product_data):
    """Parse structured product data into a flat dictionary."""
    result = {
        'name': product_data.get('name', ''),
        'description': product_data.get('description', ''),
        'sku': product_data.get('sku', ''),
        'brand': '',
        'price': '',
        'currency': '',
        'availability': '',
        'rating': '',
        'review_count': '',
    }

    # Brand
    brand = product_data.get('brand', {})
    if isinstance(brand, dict):
        result['brand'] = brand.get('name', '')
    elif isinstance(brand, str):
        result['brand'] = brand

    # Offer
    offers = product_data.get('offers', {})
    if isinstance(offers, dict):
        result['price'] = offers.get(
            'price', offers.get('lowPrice', '')
        )
        result['currency'] = offers.get('priceCurrency', '')
        result['availability'] = offers.get('availability', '')
    elif isinstance(offers, list) and offers:
        result['price'] = offers[0].get('price', '')
        result['currency'] = offers[0].get(
            'priceCurrency', ''
        )
        result['availability'] = offers[0].get(
            'availability', ''
        )

    # Rating
    rating = product_data.get('aggregateRating', {})
    if isinstance(rating, dict):
        result['rating'] = rating.get('ratingValue', '')
        result['review_count'] = rating.get(
            'reviewCount', rating.get('ratingCount', '')
        )

    return result

Common Scraping Patterns

Sitemap-based discovery

Many e-commerce sites publish sitemaps that list all product URLs:

import requests
import xml.etree.ElementTree as ET

def get_product_urls_from_sitemap(sitemap_url):
    """Extract product URLs from a sitemap."""
    response = requests.get(sitemap_url, timeout=30)
    response.raise_for_status()

    root = ET.fromstring(response.content)
    # Handle namespace
    namespace = (
        '{http://www.sitemaps.org/schemas/sitemap/0.9}'
    )

    urls = []
    for url_element in root.findall(f'{namespace}url'):
        loc = url_element.find(f'{namespace}loc')
        if loc is not None and '/product' in loc.text:
            urls.append(loc.text)

    return urls

Pagination handling

def scrape_category_pages(base_url, max_pages=50):
    """Scrape products from paginated category pages."""
    all_products = []
    page = 1

    while page <= max_pages:
        url = f"{base_url}?page={page}"
        try:
            products = extract_structured_data(url)
            if not products:
                break  # No more products

            for product in products:
                parsed = parse_product(product)
                parsed['source_url'] = url
                parsed['page'] = page
                all_products.append(parsed)

            page += 1
            import time
            time.sleep(2)  # Polite delay

        except requests.RequestException:
            break

    return all_products

Anti-Bot Challenges

Challenge How it works Ethical workaround
Rate limiting Server limits requests per IP per time period Respect rate limits; add delays between requests
CAPTCHAs Challenge-response test for humans Use APIs instead of scraping; reduce request frequency
IP blocking Server blocks IPs with unusual traffic Use ethical rate limiting; identify your bot in User-Agent
JavaScript rendering Content loaded dynamically via JavaScript Use structured data (JSON-LD) instead of parsing rendered HTML
Bot detection services Cloudflare, Akamai, PerimeterX fingerprint bots Use official APIs; structured data does not require JS rendering
Honeypot links Hidden links that trap automated crawlers Only follow visible links; respect robots.txt
Dynamic selectors CSS class names change to break scrapers Use structured data (JSON-LD) which has stable schema
Consideration Details
Terms of service Most e-commerce sites prohibit scraping in their terms of service; violating ToS may have legal consequences
Copyright Product descriptions and photos are copyrighted; collecting factual data (prices, availability) is different from copying creative content
Computer Fraud and Abuse Act (CFAA) Accessing a computer system in a way that exceeds authorised access may violate the CFAA (US law); case law is evolving
GDPR If collecting any personal data (seller names, reviewer names), GDPR may apply
robots.txt Respect robots.txt directives; they indicate which pages the site does not want crawled
Database rights (EU) The EU Database Directive protects databases; extracting substantial parts may infringe
hiQ v. LinkedIn (US) US court ruled that scraping publicly available data is not a CFAA violation, but the case was narrow and specific
Trespass to chattels Overloading a server with requests may constitute trespass to chattels

Best practices for legal compliance

Practice Details
Use official APIs when available APIs are the intended access method; terms are clear
Read and follow robots.txt Respect the site's wishes about automated access
Identify your bot Use a descriptive User-Agent string with contact information
Rate limit aggressively No more than 1 request per 2-5 seconds to any single site
Cache responses Do not re-fetch data you already have
Collect facts, not creative content Prices and availability are facts; descriptions and photos are creative works
Do not bypass authentication Do not log in to access protected data
Do not resell raw scraped data Use data for internal analysis; do not redistribute copyrighted content
Document your purpose Have a legitimate business purpose for the data you collect

Pricing Intelligence Applications

Price monitoring

Metric How to calculate Use
Price index Your price / average competitor price * 100 Are you above or below market?
Price position Rank among competitors by price Where you sit in the market
Price change frequency Number of price changes per product per month How dynamic is competitor pricing?
Price elasticity Change in sales rank / change in price How sensitive is demand to price?
MAP violations Count of resellers below MAP Enforce pricing agreements
Promotional frequency Number of sale events per product per month Competitor promotional strategy
Stock-out rate Percentage of time a product shows "out of stock" Supply chain health

Python: Simple price tracker

import csv
import os
from datetime import datetime

def track_price(product_id, product_name, price,
                currency, source, output_file):
    """Append a price observation to a tracking file."""
    file_exists = os.path.exists(output_file)

    with open(output_file, 'a', newline='') as f:
        writer = csv.writer(f)
        if not file_exists:
            writer.writerow([
                'timestamp', 'product_id', 'product_name',
                'price', 'currency', 'source'
            ])
        writer.writerow([
            datetime.now().isoformat(),
            product_id,
            product_name,
            price,
            currency,
            source
        ])

Extracting Seller Contact Information

E-commerce marketplace pages sometimes display seller contact information (email addresses, phone numbers, website links). When collecting seller data for B2B outreach purposes, save product or seller pages as HTML files and upload to Email Extractor to extract email addresses. This is useful for building outreach lists of marketplace sellers, brand owners or suppliers.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)