Scraping Pricing Data: Competitive Intelligence and Market Monitoring
On this page
Why Scrape Pricing Data
Pricing is one of the most valuable and time-sensitive forms of competitive intelligence. Companies scrape pricing data to:
- Monitor competitors. Track competitor prices in real time to respond quickly to changes.
- Optimise pricing. Set prices based on market positioning rather than guesswork.
- Detect MAP violations. Monitor minimum advertised price compliance across reseller channels.
- Track market trends. Identify pricing trends by category, brand or geography.
- Inform sourcing. Compare supplier pricing across regions and platforms.
- Support sales. Arm sales teams with competitive pricing intelligence for negotiations.
- Feed dynamic pricing. Provide real-time market data to algorithmic pricing engines.
Data Sources
E-commerce marketplaces
| Platform | Data available | Access difficulty |
|---|---|---|
| Amazon | Price, buy box, seller, reviews, ratings, stock | High (aggressive anti-scraping) |
| Walmart | Price, availability, seller, ratings | High |
| eBay | Price, bids, seller, condition, shipping | Medium (API available) |
| Etsy | Price, seller, reviews, shipping | Medium |
| AliExpress/Alibaba | Price, MOQ, seller, shipping, variations | Medium |
| Shopify stores | Price, variants, stock (via /products.json) | Low (many expose product API) |
SaaS and subscription pricing
| Source | Data available | Access difficulty |
|---|---|---|
| Public pricing pages | Plan names, features, prices | Low |
| Comparison sites (G2, Capterra) | Pricing tiers, user reviews | Medium |
| API documentation | API pricing, rate limits, usage tiers | Low |
| Archived pricing pages | Historical pricing changes | Low (Wayback Machine) |
B2B and wholesale
| Source | Data available | Access difficulty |
|---|---|---|
| Distributor portals | Wholesale pricing, stock levels | High (authenticated access) |
| Government procurement sites | Contract pricing, RFP responses | Low (public records) |
| Industry benchmarks | Market rate data, salary data, material costs | Varies |
| Trade publications | Market pricing reports, commodity indices | Medium |
Travel and hospitality
| Source | Data available | Access difficulty |
|---|---|---|
| OTAs (Booking, Expedia) | Room rates, availability, property details | High (aggressive anti-scraping) |
| Airline sites | Fare prices, route availability | Very high |
| Google Flights/Hotels | Aggregated pricing | Very high |
| Metasearch (Kayak, Trivago) | Comparison pricing | High |
Scraping Techniques
HTML parsing
Most pricing data is embedded in HTML product pages. Extract with:
CSS selectors (Python with Beautiful Soup):
import requests
from bs4 import BeautifulSoup
response = requests.get(
"https://example.com/product/widget",
headers={"User-Agent": "Mozilla/5.0"}
)
soup = BeautifulSoup(response.text, "html.parser")
price_element = soup.select_one(".product-price .current-price")
if price_element:
price_text = price_element.get_text(strip=True)
# Clean the price: remove currency symbol, commas
price = float(price_text.replace("$", "").replace(",", ""))
XPath (Python with lxml):
from lxml import html
import requests
response = requests.get(
"https://example.com/product/widget",
headers={"User-Agent": "Mozilla/5.0"}
)
tree = html.fromstring(response.content)
price = tree.xpath('//span[@class="price"]/text()')
Structured data extraction
Many e-commerce sites include structured data (JSON-LD or Microdata) with pricing information:
import json
from bs4 import BeautifulSoup
soup = BeautifulSoup(response.text, "html.parser")
for script in soup.find_all("script", type="application/ld+json"):
data = json.loads(script.string)
if data.get("@type") == "Product":
offers = data.get("offers", {})
if isinstance(offers, dict):
price = offers.get("price")
currency = offers.get("priceCurrency")
availability = offers.get("availability")
Structured data is the cleanest source because:
- It follows a standard schema (Schema.org).
- It is machine-readable by design.
- It is less likely to change layout than visual HTML.
- It often includes availability, condition and seller information.
See Structured Data Extraction Guide for detailed coverage.
JavaScript-rendered pricing
Some sites load pricing dynamically with JavaScript (React, Vue, Angular). Static HTML parsing will not find the price.
Solutions:
| Approach | Tool | When to use |
|---|---|---|
| Headless browser | Playwright, Puppeteer, Selenium | Price loaded by JavaScript, requires page rendering |
| API interception | Browser DevTools, mitmproxy | Price loaded from a separate API call |
| Direct API call | requests, httpx | You've identified the API endpoint that returns pricing |
Playwright example:
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com/product/widget")
page.wait_for_selector(".product-price")
price_text = page.inner_text(".product-price")
browser.close()
API-based extraction
Some platforms offer APIs with pricing data:
| Platform | API | Pricing data available |
|---|---|---|
| eBay | Browse API | Current listings, prices, bids |
| Amazon | Product Advertising API | Prices (for affiliates) |
| Shopify | Storefront API, /products.json | Public product and variant prices |
| Best Buy | Products API | Prices, availability, specs |
| Walmart | Affiliate API | Prices, availability |
| Google Shopping | Content API for Shopping | Price benchmarks (for merchants) |
APIs are always preferred over scraping because:
- Access is explicitly permitted.
- Data is structured and consistent.
- Rate limits are documented.
- No risk of being blocked.
Shopify store shortcut
Many Shopify stores expose product data at /products.json:
import requests
response = requests.get("https://example-store.myshopify.com/products.json")
data = response.json()
for product in data["products"]:
title = product["title"]
for variant in product["variants"]:
price = variant["price"]
compare_at = variant.get("compare_at_price")
sku = variant.get("sku")
available = variant.get("available")
This endpoint returns up to 250 products per page. Paginate with ?page=2, ?page=3, etc.
Building a Price Monitoring System
Architecture
A price monitoring system has five components:
- URL management. A database of product URLs to monitor.
- Scraper. Extracts pricing data from each URL on a schedule.
- Data storage. Stores historical pricing data.
- Change detection. Identifies price changes and triggers alerts.
- Reporting. Dashboards and reports for analysis.
URL management
Maintain a database of product URLs to monitor:
| Field | Purpose |
|---|---|
| URL | The product page to scrape |
| Competitor | Which competitor this product belongs to |
| Product category | For grouping and analysis |
| Your SKU | Your equivalent product (for comparison) |
| Your price | Your current price (for competitive comparison) |
| Scrape frequency | How often to check (hourly, daily, weekly) |
| Last scraped | Timestamp of last successful scrape |
| Status | Active, paused, error |
Scheduling
| Pricing type | Recommended frequency |
|---|---|
| E-commerce (high competition) | Every 1-4 hours |
| E-commerce (moderate competition) | Daily |
| SaaS pricing pages | Weekly |
| B2B / wholesale | Weekly or monthly |
| Travel / hospitality | Hourly (prices change frequently) |
| Commodities | Real-time or hourly |
Change detection and alerts
Alert types:
| Alert | Trigger | Action |
|---|---|---|
| Price drop | Competitor price drops below your price | Review and consider matching |
| Price increase | Competitor price increases above your price | Opportunity to increase margin |
| New product | New product detected at a competitor | Analyse and respond |
| Out of stock | Competitor product goes out of stock | Opportunity to capture demand |
| Back in stock | Competitor product returns to stock | Monitor for price changes |
| Significant change | Price changes by more than X% | Review for errors or strategy shifts |
Data storage
Store historical pricing data for trend analysis:
| Field | Type | Purpose |
|---|---|---|
| url | string | Product page URL |
| scraped_at | timestamp | When the data was collected |
| price | decimal | The scraped price |
| currency | string | Currency code |
| availability | boolean | In stock or not |
| seller | string | Who is selling (for marketplace products) |
| shipping | decimal | Shipping cost if shown |
| original_price | decimal | "Was" price or list price |
| discount_pct | decimal | Calculated discount percentage |
No-Code and Commercial Tools
| Tool | What it does | Pricing |
|---|---|---|
| Prisync | Competitor price monitoring for e-commerce | From $99/month |
| Competera | AI-driven pricing optimisation | Enterprise pricing |
| Price2Spy | Price monitoring and MAP compliance | From $24/month |
| Skuuudle | Competitor price tracking | Custom pricing |
| Intelligence Node | Retail analytics and price intelligence | Enterprise pricing |
| Keepa | Amazon price history tracking | Free (browser extension) + paid API |
| CamelCamelCamel | Amazon price history | Free |
| Scrapy + Splash | Open-source scraping framework | Free (self-hosted) |
| Apify | Cloud scraping platform with pre-built scrapers | From $49/month |
| Bright Data | Proxy network + scraping infrastructure | From $500/month |
Data Analysis
Price position analysis
Calculate your price position relative to competitors:
| Metric | Calculation | Interpretation |
|---|---|---|
| Price index | (Your price / Competitor average price) x 100 | 100 = at market; below 100 = below market; above 100 = above market |
| Price gap | Your price - Competitor price | Absolute difference |
| Price rank | Your position among all sellers | 1st = cheapest |
| Promotional frequency | Percentage of time on promotion | How often competitors discount |
| Promotional depth | Average discount when on promotion | How deeply competitors discount |
Trend analysis
Track pricing trends over time:
- Seasonal patterns (when do competitors raise/lower prices?).
- Response patterns (how quickly do competitors match your price changes?).
- Category trends (are prices rising or falling across the category?).
- New entrant impact (how did a new competitor affect market pricing?).
Reporting
| Report | Frequency | Audience |
|---|---|---|
| Daily price change summary | Daily | Pricing team, merchandising |
| Competitive price position | Weekly | Category managers, marketing |
| Market trend analysis | Monthly | Leadership, strategy |
| MAP compliance report | Weekly | Brand management, legal |
| Price elasticity analysis | Quarterly | Pricing strategy, finance |
Legal and Ethical Considerations
Legal landscape
- Terms of service. Most e-commerce sites prohibit scraping in their ToS.
- CFAA (US). Accessing a computer system without authorisation is a federal crime. However, scraping publicly available pricing data has generally been found not to violate the CFAA (hiQ v. LinkedIn).
- Copyright. Individual prices are not copyrightable. However, a compiled database of pricing data may have copyright protection (especially in the EU under the Database Directive).
- Tortious interference. Aggressive scraping that harms a competitor's website (slowing it down, increasing their costs) could create tort liability.
Ethical practices
- Rate limit requests to avoid impacting the target site's performance.
- Respect robots.txt (although compliance is voluntary, not legally required in most jurisdictions).
- Do not circumvent authentication or access controls.
- Do not scrape personal data (pricing data is generally not personal data).
- Do not misrepresent yourself (do not fake login credentials).
- Use the data for internal competitive intelligence, not for republishing or resale (unless you have the right to do so).
- Prefer APIs and public data sources over scraping when available.
See Is Web Scraping Legal and Scraping Ethics Best Practices.
Integrating Contact Data
Pricing intelligence often leads to outreach: contacting suppliers, distributors or potential partners based on pricing data.
After scraping pricing data, you may export contact information (seller names, company domains, support emails) alongside the pricing data. Upload these exported files to Email Extractor to extract and deduplicate email addresses from the scraped data.