PA API (affiliate programme); Amazon SP-API (sellers)
Shopify stores
Storefront API (if enabled by merchant)
Varies by store
Structured data (JSON-LD); Storefront API where available
eBay
Browse API, Finding API
Restricted in terms of service
eBay APIs (developer programme)
Walmart
Affiliate API
Prohibited in terms of service
Walmart API (affiliate programme)
Etsy
Open API
Restricted
Etsy API (developer programme)
Best Buy
Products API
Restricted
Best Buy API (developer programme)
Target
No public API
Prohibited
No official access
AliExpress
Affiliate API
Restricted
AliExpress API (affiliate)
Structured data on product pages
Many e-commerce sites include structured data (JSON-LD, Microdata, RDFa) in their HTML that is designed for search engines. This data is machine-readable and often includes:
Schema.org type
Fields typically included
Where to find
Product
Name, description, SKU, brand, image, category
JSON-LD in page <head> or <body>
Offer
Price, currency, availability, seller, condition
Nested within Product
AggregateRating
Rating value, review count
Nested within Product
Review
Author, rating, review body, date
Nested within Product or separate
BreadcrumbList
Category hierarchy
JSON-LD in page <head>
Extracting structured data with Python
import requests
from bs4 import BeautifulSoup
import json
def extract_structured_data(url, headers=None):
"""Extract JSON-LD structured data from a product page."""
if headers is None:
headers = {
'User-Agent': (
'Mozilla/5.0 (compatible; '
'DataResearchBot/1.0; '
'+https://example.com/bot)'
)
}
response = requests.get(url, headers=headers, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')
structured_data = []
# Find all JSON-LD blocks
for script in soup.find_all(
'script', type='application/ld+json'
):
try:
data = json.loads(script.string)
structured_data.append(data)
except (json.JSONDecodeError, TypeError):
continue
# Filter for product data
products = []
for data in structured_data:
if isinstance(data, list):
for item in data:
if item.get('@type') == 'Product':
products.append(item)
elif isinstance(data, dict):
if data.get('@type') == 'Product':
products.append(data)
# Check @graph
if '@graph' in data:
for item in data['@graph']:
if item.get('@type') == 'Product':
products.append(item)
return products
def parse_product(product_data):
"""Parse structured product data into a flat dictionary."""
result = {
'name': product_data.get('name', ''),
'description': product_data.get('description', ''),
'sku': product_data.get('sku', ''),
'brand': '',
'price': '',
'currency': '',
'availability': '',
'rating': '',
'review_count': '',
}
# Brand
brand = product_data.get('brand', {})
if isinstance(brand, dict):
result['brand'] = brand.get('name', '')
elif isinstance(brand, str):
result['brand'] = brand
# Offer
offers = product_data.get('offers', {})
if isinstance(offers, dict):
result['price'] = offers.get(
'price', offers.get('lowPrice', '')
)
result['currency'] = offers.get('priceCurrency', '')
result['availability'] = offers.get('availability', '')
elif isinstance(offers, list) and offers:
result['price'] = offers[0].get('price', '')
result['currency'] = offers[0].get(
'priceCurrency', ''
)
result['availability'] = offers[0].get(
'availability', ''
)
# Rating
rating = product_data.get('aggregateRating', {})
if isinstance(rating, dict):
result['rating'] = rating.get('ratingValue', '')
result['review_count'] = rating.get(
'reviewCount', rating.get('ratingCount', '')
)
return result
Common Scraping Patterns
Sitemap-based discovery
Many e-commerce sites publish sitemaps that list all product URLs:
import requests
import xml.etree.ElementTree as ET
def get_product_urls_from_sitemap(sitemap_url):
"""Extract product URLs from a sitemap."""
response = requests.get(sitemap_url, timeout=30)
response.raise_for_status()
root = ET.fromstring(response.content)
# Handle namespace
namespace = (
'{http://www.sitemaps.org/schemas/sitemap/0.9}'
)
urls = []
for url_element in root.findall(f'{namespace}url'):
loc = url_element.find(f'{namespace}loc')
if loc is not None and '/product' in loc.text:
urls.append(loc.text)
return urls
Pagination handling
def scrape_category_pages(base_url, max_pages=50):
"""Scrape products from paginated category pages."""
all_products = []
page = 1
while page <= max_pages:
url = f"{base_url}?page={page}"
try:
products = extract_structured_data(url)
if not products:
break # No more products
for product in products:
parsed = parse_product(product)
parsed['source_url'] = url
parsed['page'] = page
all_products.append(parsed)
page += 1
import time
time.sleep(2) # Polite delay
except requests.RequestException:
break
return all_products
Anti-Bot Challenges
Challenge
How it works
Ethical workaround
Rate limiting
Server limits requests per IP per time period
Respect rate limits; add delays between requests
CAPTCHAs
Challenge-response test for humans
Use APIs instead of scraping; reduce request frequency
IP blocking
Server blocks IPs with unusual traffic
Use ethical rate limiting; identify your bot in User-Agent
JavaScript rendering
Content loaded dynamically via JavaScript
Use structured data (JSON-LD) instead of parsing rendered HTML
Bot detection services
Cloudflare, Akamai, PerimeterX fingerprint bots
Use official APIs; structured data does not require JS rendering
Honeypot links
Hidden links that trap automated crawlers
Only follow visible links; respect robots.txt
Dynamic selectors
CSS class names change to break scrapers
Use structured data (JSON-LD) which has stable schema
Legal Considerations
Consideration
Details
Terms of service
Most e-commerce sites prohibit scraping in their terms of service; violating ToS may have legal consequences
Copyright
Product descriptions and photos are copyrighted; collecting factual data (prices, availability) is different from copying creative content
Computer Fraud and Abuse Act (CFAA)
Accessing a computer system in a way that exceeds authorised access may violate the CFAA (US law); case law is evolving
GDPR
If collecting any personal data (seller names, reviewer names), GDPR may apply
robots.txt
Respect robots.txt directives; they indicate which pages the site does not want crawled
Database rights (EU)
The EU Database Directive protects databases; extracting substantial parts may infringe
hiQ v. LinkedIn (US)
US court ruled that scraping publicly available data is not a CFAA violation, but the case was narrow and specific
Trespass to chattels
Overloading a server with requests may constitute trespass to chattels
Best practices for legal compliance
Practice
Details
Use official APIs when available
APIs are the intended access method; terms are clear
Read and follow robots.txt
Respect the site's wishes about automated access
Identify your bot
Use a descriptive User-Agent string with contact information
Rate limit aggressively
No more than 1 request per 2-5 seconds to any single site
Cache responses
Do not re-fetch data you already have
Collect facts, not creative content
Prices and availability are facts; descriptions and photos are creative works
Do not bypass authentication
Do not log in to access protected data
Do not resell raw scraped data
Use data for internal analysis; do not redistribute copyrighted content
Document your purpose
Have a legitimate business purpose for the data you collect
Pricing Intelligence Applications
Price monitoring
Metric
How to calculate
Use
Price index
Your price / average competitor price * 100
Are you above or below market?
Price position
Rank among competitors by price
Where you sit in the market
Price change frequency
Number of price changes per product per month
How dynamic is competitor pricing?
Price elasticity
Change in sales rank / change in price
How sensitive is demand to price?
MAP violations
Count of resellers below MAP
Enforce pricing agreements
Promotional frequency
Number of sale events per product per month
Competitor promotional strategy
Stock-out rate
Percentage of time a product shows "out of stock"
Supply chain health
Python: Simple price tracker
import csv
import os
from datetime import datetime
def track_price(product_id, product_name, price,
currency, source, output_file):
"""Append a price observation to a tracking file."""
file_exists = os.path.exists(output_file)
with open(output_file, 'a', newline='') as f:
writer = csv.writer(f)
if not file_exists:
writer.writerow([
'timestamp', 'product_id', 'product_name',
'price', 'currency', 'source'
])
writer.writerow([
datetime.now().isoformat(),
product_id,
product_name,
price,
currency,
source
])
Extracting Seller Contact Information
E-commerce marketplace pages sometimes display seller contact information (email addresses, phone numbers, website links). When collecting seller data for B2B outreach purposes, save product or seller pages as HTML files and upload to Email Extractor to extract email addresses. This is useful for building outreach lists of marketplace sellers, brand owners or suppliers.