Structured Data Extraction: Turning Unstructured Web Pages into Usable Datasets
On this page
What Structured Data Extraction Means
Structured data extraction is the process of turning unstructured or semi-structured content (web pages, PDFs, emails, documents) into clean, organised datasets that can be analysed, stored in a database or fed into another system.
A web page is semi-structured. It has HTML markup that gives it some structure, but the data you want (a price, a name, an address, an email, a product specification) is mixed in with navigation, advertising, styling and layout elements. Structured data extraction separates the signal from the noise.
Examples of structured data extraction:
| Source | What you extract | Output format |
|---|---|---|
| Product pages | Name, price, description, SKU, availability | CSV or database table |
| Job postings | Title, company, location, salary, requirements | JSON or spreadsheet |
| Directory listings | Business name, address, phone, email, category | CSV |
| Real estate listings | Address, price, bedrooms, bathrooms, square footage | Database |
| Academic papers | Title, authors, abstract, citations, DOI | BibTeX or JSON |
| News articles | Headline, author, date, body text, category | JSON |
| Email files | Sender, recipient, subject, date, email addresses | CSV |
For email address extraction specifically, upload your files to Email Extractor, which pulls email addresses from 19 file formats including HTML, XML, JSON, CSV, PDF, DOCX, XLSX and more.
How Web Pages Store Data
HTML structure
HTML is a tree of nested elements. Data lives inside elements, as text content or as attribute values.
<div class="product-card">
<h2 class="product-name">Widget Pro 3000</h2>
<span class="price">$49.99</span>
<p class="description">The best widget for professional use.</p>
<a href="/products/widget-pro-3000" class="product-link">View details</a>
<span class="availability in-stock">In Stock</span>
</div>
In this example:
- The product name is the text inside the
<h2>with classproduct-name. - The price is the text inside the
<span>with classprice. - The product URL is the
hrefattribute of the<a>tag. - Availability is both the text content and the class name of the
<span>.
Structured data markup
Many websites embed structured data in machine-readable formats alongside the human-readable HTML. This makes extraction much easier when it is present.
JSON-LD (JavaScript Object Notation for Linked Data):
The most common format. Embedded in a <script> tag:
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Product",
"name": "Widget Pro 3000",
"description": "The best widget for professional use.",
"offers": {
"@type": "Offer",
"price": "49.99",
"priceCurrency": "USD",
"availability": "https://schema.org/InStock"
}
}
</script>
Microdata:
Embedded as attributes on HTML elements:
<div itemscope itemtype="https://schema.org/Product">
<h2 itemprop="name">Widget Pro 3000</h2>
<span itemprop="price" content="49.99">$49.99</span>
</div>
RDFa:
Similar to microdata but uses different attributes:
<div vocab="https://schema.org/" typeof="Product">
<h2 property="name">Widget Pro 3000</h2>
<span property="price" content="49.99">$49.99</span>
</div>
APIs
Some websites provide APIs that return structured data directly, making scraping unnecessary:
- Public APIs. Many sites offer free or paid APIs for their data (GitHub, Twitter/X, Reddit, many e-commerce platforms).
- Internal APIs. JavaScript-heavy websites often fetch data from internal API endpoints. Browser developer tools (Network tab) reveal these endpoints, and they often return clean JSON.
- RSS feeds. Blogs and news sites offer RSS feeds with structured article data.
Always check for an API before building a scraper. API data is cleaner, more stable and explicitly permitted.
Extraction Methods
CSS selectors
CSS selectors identify HTML elements by tag name, class, ID, attributes and position in the document tree.
| Selector | What it matches | Example |
|---|---|---|
div |
All <div> elements |
Tag name |
.price |
Elements with class "price" | Class |
#main-content |
Element with ID "main-content" | ID |
div.product-card |
<div> elements with class "product-card" |
Tag + class |
div > h2 |
<h2> directly inside a <div> |
Child combinator |
div h2 |
<h2> anywhere inside a <div> |
Descendant combinator |
a[href] |
<a> elements with an href attribute |
Attribute presence |
a[href^="https"] |
<a> with href starting with "https" |
Attribute prefix |
span:nth-child(2) |
The second <span> child |
Position |
Python example with Beautiful Soup:
from bs4 import BeautifulSoup
html = '<div class="product-card"><h2 class="product-name">Widget</h2><span class="price">$49.99</span></div>'
soup = BeautifulSoup(html, 'html.parser')
name = soup.select_one('.product-name').text
price = soup.select_one('.price').text
XPath
XPath is more powerful than CSS selectors. It can navigate up the tree (to parents), use conditions, and match text content.
| XPath | What it matches |
|---|---|
//div[@class="product-card"] |
All <div> with class "product-card" |
//h2[contains(@class, "name")] |
All <h2> with "name" in the class |
//span[@class="price"]/text() |
Text content of <span class="price"> |
//a[contains(text(), "View")]/@href |
href of <a> tags containing "View" |
//div[@class="product-card"]//span |
All <span> inside product cards |
//table/tr[position()>1]/td[2] |
Second column of all rows except the header |
Python example with lxml:
from lxml import html
tree = html.fromstring(page_content)
names = tree.xpath('//h2[@class="product-name"]/text()')
prices = tree.xpath('//span[@class="price"]/text()')
Regular expressions
Regular expressions match text patterns. They are useful for extracting data from unstructured text (not HTML).
| Pattern | What it matches |
|---|---|
[\w.-]+@[\w.-]+\.\w+ |
Email addresses |
\$[\d,]+\.?\d* |
Dollar amounts |
\(\d{3}\)\s?\d{3}-\d{4} |
US phone numbers (xxx) xxx-xxxx |
\d{5}(-\d{4})? |
US ZIP codes |
https?://[\w./%-]+ |
URLs |
Warning: Do not parse HTML with regular expressions. HTML is a nested, recursive structure that regular expressions cannot handle correctly. Use a proper HTML parser (Beautiful Soup, lxml) for HTML, and regular expressions only for flat text patterns within extracted text.
JSON-LD extraction
When a page contains JSON-LD, extraction is trivial:
import json
from bs4 import BeautifulSoup
soup = BeautifulSoup(page_content, 'html.parser')
scripts = soup.find_all('script', type='application/ld+json')
for script in scripts:
data = json.loads(script.string)
if data.get('@type') == 'Product':
print(data['name'], data['offers']['price'])
JSON-LD is pre-structured. If a page has it, always extract from JSON-LD rather than parsing the HTML.
Building an Extraction Pipeline
Step 1: Analyse the source
Before writing code, understand the source:
- Open the page in a browser. Look at the data you want to extract.
- View the page source. Check for JSON-LD or structured data markup.
- Inspect elements. Use browser developer tools (right-click, Inspect) to examine the HTML structure around the data.
- Check the Network tab. Look for API calls that return the data as JSON.
- Check robots.txt. Review the site's robots.txt for scraping policies.
- Look for an API. Search for "[site name] API" before building a scraper.
Step 2: Build the extractor
For a single page type (e.g., all product pages on one site):
import requests
from bs4 import BeautifulSoup
import csv
import time
def extract_product(url):
response = requests.get(url, headers={
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64)'
})
soup = BeautifulSoup(response.text, 'html.parser')
# Try JSON-LD first
ld_scripts = soup.find_all('script', type='application/ld+json')
for script in ld_scripts:
try:
data = json.loads(script.string)
if data.get('@type') == 'Product':
return {
'name': data.get('name'),
'price': data.get('offers', {}).get('price'),
'description': data.get('description'),
'url': url
}
except json.JSONDecodeError:
continue
# Fall back to HTML parsing
return {
'name': soup.select_one('.product-name').text.strip() if soup.select_one('.product-name') else None,
'price': soup.select_one('.price').text.strip() if soup.select_one('.price') else None,
'description': soup.select_one('.description').text.strip() if soup.select_one('.description') else None,
'url': url
}
Step 3: Handle pagination and listing pages
Most scraping targets have listing pages (search results, category pages) that link to detail pages.
def get_listing_urls(listing_url):
"""Extract product URLs from a listing page."""
response = requests.get(listing_url)
soup = BeautifulSoup(response.text, 'html.parser')
product_urls = []
for link in soup.select('.product-card a.product-link'):
href = link.get('href')
if href:
full_url = urljoin(listing_url, href)
product_urls.append(full_url)
# Handle pagination
next_page = soup.select_one('a.next-page')
if next_page:
next_url = urljoin(listing_url, next_page['href'])
product_urls.extend(get_listing_urls(next_url))
return product_urls
Step 4: Clean the data
Raw extracted data is messy. Clean it before storage:
| Problem | Example | Fix |
|---|---|---|
| Extra whitespace | " Widget Pro 3000 " | .strip(), collapse internal spaces |
| HTML entities | "Widget & Gadget" | html.unescape() |
| Currency symbols in numbers | "$49.99" | Remove $, convert to float |
| Inconsistent formats | "49.99", "49,99", "$49.99" | Normalise to one format |
| Missing values | None, empty string, "N/A" |
Standardise to None or empty |
| Unicode issues | Mixed encodings, mojibake | Detect and convert to UTF-8 |
Step 5: Store the data
| Format | When to use |
|---|---|
| CSV | Small to medium datasets, spreadsheet-compatible, human-readable |
| JSON | Nested data, API consumption, flexible schema |
| SQLite | Structured queries, deduplication, relationships |
| PostgreSQL/MySQL | Production use, multiple users, large datasets |
| Parquet | Large analytical datasets, column-oriented queries |
Tools
Libraries
| Tool | Language | Best for |
|---|---|---|
| Beautiful Soup | Python | HTML parsing, simple scraping |
| lxml | Python | Fast HTML/XML parsing, XPath |
| Scrapy | Python | Large-scale scraping with built-in pipeline |
| Playwright | Python/JS | JavaScript-rendered pages |
| Selenium | Python/JS/Java | Browser automation, dynamic pages |
| Cheerio | JavaScript | HTML parsing in Node.js |
| Puppeteer | JavaScript | Headless Chrome automation |
Managed services
| Service | What it does |
|---|---|
| Apify | Cloud scraping platform with pre-built actors |
| Bright Data | Proxy network + scraping tools |
| ScrapingBee | API for scraping with JavaScript rendering |
| Diffbot | AI-powered structured data extraction |
| Import.io | Visual web data extraction |
No-code tools
| Tool | How it works |
|---|---|
| Octoparse | Visual point-and-click scraper |
| Browse AI | Train a robot by showing it what to extract |
| ParseHub | Visual scraping with pagination handling |
| Web Scraper (Chrome extension) | Build scrapers in the browser |
Legal and Ethical Considerations
- Robots.txt. Check and respect the site's robots.txt file.
- Terms of service. Many sites prohibit scraping in their ToS.
- Rate limiting. Do not overload servers. Add delays between requests.
- Copyright. Extracted data may be copyrighted. Extraction for analysis and personal use is generally more defensible than republication.
- Personal data. Extracting and storing personal data must comply with GDPR, CCPA and other privacy regulations.
See Is Web Scraping Legal and Scraping Ethics Best Practices.