Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Structured Data Extraction: Turning Unstructured Web Pages into Usable Datasets

On this page

What Structured Data Extraction Means

Structured data extraction is the process of turning unstructured or semi-structured content (web pages, PDFs, emails, documents) into clean, organised datasets that can be analysed, stored in a database or fed into another system.

A web page is semi-structured. It has HTML markup that gives it some structure, but the data you want (a price, a name, an address, an email, a product specification) is mixed in with navigation, advertising, styling and layout elements. Structured data extraction separates the signal from the noise.

Examples of structured data extraction:

Source What you extract Output format
Product pages Name, price, description, SKU, availability CSV or database table
Job postings Title, company, location, salary, requirements JSON or spreadsheet
Directory listings Business name, address, phone, email, category CSV
Real estate listings Address, price, bedrooms, bathrooms, square footage Database
Academic papers Title, authors, abstract, citations, DOI BibTeX or JSON
News articles Headline, author, date, body text, category JSON
Email files Sender, recipient, subject, date, email addresses CSV

For email address extraction specifically, upload your files to Email Extractor, which pulls email addresses from 19 file formats including HTML, XML, JSON, CSV, PDF, DOCX, XLSX and more.

How Web Pages Store Data

HTML structure

HTML is a tree of nested elements. Data lives inside elements, as text content or as attribute values.

<div class="product-card">
  <h2 class="product-name">Widget Pro 3000</h2>
  <span class="price">$49.99</span>
  <p class="description">The best widget for professional use.</p>
  <a href="/products/widget-pro-3000" class="product-link">View details</a>
  <span class="availability in-stock">In Stock</span>
</div>

In this example:

  • The product name is the text inside the <h2> with class product-name.
  • The price is the text inside the <span> with class price.
  • The product URL is the href attribute of the <a> tag.
  • Availability is both the text content and the class name of the <span>.

Structured data markup

Many websites embed structured data in machine-readable formats alongside the human-readable HTML. This makes extraction much easier when it is present.

JSON-LD (JavaScript Object Notation for Linked Data):

The most common format. Embedded in a <script> tag:

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Product",
  "name": "Widget Pro 3000",
  "description": "The best widget for professional use.",
  "offers": {
    "@type": "Offer",
    "price": "49.99",
    "priceCurrency": "USD",
    "availability": "https://schema.org/InStock"
  }
}
</script>

Microdata:

Embedded as attributes on HTML elements:

<div itemscope itemtype="https://schema.org/Product">
  <h2 itemprop="name">Widget Pro 3000</h2>
  <span itemprop="price" content="49.99">$49.99</span>
</div>

RDFa:

Similar to microdata but uses different attributes:

<div vocab="https://schema.org/" typeof="Product">
  <h2 property="name">Widget Pro 3000</h2>
  <span property="price" content="49.99">$49.99</span>
</div>

APIs

Some websites provide APIs that return structured data directly, making scraping unnecessary:

  • Public APIs. Many sites offer free or paid APIs for their data (GitHub, Twitter/X, Reddit, many e-commerce platforms).
  • Internal APIs. JavaScript-heavy websites often fetch data from internal API endpoints. Browser developer tools (Network tab) reveal these endpoints, and they often return clean JSON.
  • RSS feeds. Blogs and news sites offer RSS feeds with structured article data.

Always check for an API before building a scraper. API data is cleaner, more stable and explicitly permitted.

Extraction Methods

CSS selectors

CSS selectors identify HTML elements by tag name, class, ID, attributes and position in the document tree.

Selector What it matches Example
div All <div> elements Tag name
.price Elements with class "price" Class
#main-content Element with ID "main-content" ID
div.product-card <div> elements with class "product-card" Tag + class
div > h2 <h2> directly inside a <div> Child combinator
div h2 <h2> anywhere inside a <div> Descendant combinator
a[href] <a> elements with an href attribute Attribute presence
a[href^="https"] <a> with href starting with "https" Attribute prefix
span:nth-child(2) The second <span> child Position

Python example with Beautiful Soup:

from bs4 import BeautifulSoup

html = '<div class="product-card"><h2 class="product-name">Widget</h2><span class="price">$49.99</span></div>'
soup = BeautifulSoup(html, 'html.parser')

name = soup.select_one('.product-name').text
price = soup.select_one('.price').text

XPath

XPath is more powerful than CSS selectors. It can navigate up the tree (to parents), use conditions, and match text content.

XPath What it matches
//div[@class="product-card"] All <div> with class "product-card"
//h2[contains(@class, "name")] All <h2> with "name" in the class
//span[@class="price"]/text() Text content of <span class="price">
//a[contains(text(), "View")]/@href href of <a> tags containing "View"
//div[@class="product-card"]//span All <span> inside product cards
//table/tr[position()>1]/td[2] Second column of all rows except the header

Python example with lxml:

from lxml import html

tree = html.fromstring(page_content)
names = tree.xpath('//h2[@class="product-name"]/text()')
prices = tree.xpath('//span[@class="price"]/text()')

Regular expressions

Regular expressions match text patterns. They are useful for extracting data from unstructured text (not HTML).

Pattern What it matches
[\w.-]+@[\w.-]+\.\w+ Email addresses
\$[\d,]+\.?\d* Dollar amounts
\(\d{3}\)\s?\d{3}-\d{4} US phone numbers (xxx) xxx-xxxx
\d{5}(-\d{4})? US ZIP codes
https?://[\w./%-]+ URLs

Warning: Do not parse HTML with regular expressions. HTML is a nested, recursive structure that regular expressions cannot handle correctly. Use a proper HTML parser (Beautiful Soup, lxml) for HTML, and regular expressions only for flat text patterns within extracted text.

JSON-LD extraction

When a page contains JSON-LD, extraction is trivial:

import json
from bs4 import BeautifulSoup

soup = BeautifulSoup(page_content, 'html.parser')
scripts = soup.find_all('script', type='application/ld+json')

for script in scripts:
    data = json.loads(script.string)
    if data.get('@type') == 'Product':
        print(data['name'], data['offers']['price'])

JSON-LD is pre-structured. If a page has it, always extract from JSON-LD rather than parsing the HTML.

Building an Extraction Pipeline

Step 1: Analyse the source

Before writing code, understand the source:

  1. Open the page in a browser. Look at the data you want to extract.
  2. View the page source. Check for JSON-LD or structured data markup.
  3. Inspect elements. Use browser developer tools (right-click, Inspect) to examine the HTML structure around the data.
  4. Check the Network tab. Look for API calls that return the data as JSON.
  5. Check robots.txt. Review the site's robots.txt for scraping policies.
  6. Look for an API. Search for "[site name] API" before building a scraper.

Step 2: Build the extractor

For a single page type (e.g., all product pages on one site):

import requests
from bs4 import BeautifulSoup
import csv
import time

def extract_product(url):
    response = requests.get(url, headers={
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64)'
    })
    soup = BeautifulSoup(response.text, 'html.parser')

    # Try JSON-LD first
    ld_scripts = soup.find_all('script', type='application/ld+json')
    for script in ld_scripts:
        try:
            data = json.loads(script.string)
            if data.get('@type') == 'Product':
                return {
                    'name': data.get('name'),
                    'price': data.get('offers', {}).get('price'),
                    'description': data.get('description'),
                    'url': url
                }
        except json.JSONDecodeError:
            continue

    # Fall back to HTML parsing
    return {
        'name': soup.select_one('.product-name').text.strip() if soup.select_one('.product-name') else None,
        'price': soup.select_one('.price').text.strip() if soup.select_one('.price') else None,
        'description': soup.select_one('.description').text.strip() if soup.select_one('.description') else None,
        'url': url
    }

Step 3: Handle pagination and listing pages

Most scraping targets have listing pages (search results, category pages) that link to detail pages.

def get_listing_urls(listing_url):
    """Extract product URLs from a listing page."""
    response = requests.get(listing_url)
    soup = BeautifulSoup(response.text, 'html.parser')

    product_urls = []
    for link in soup.select('.product-card a.product-link'):
        href = link.get('href')
        if href:
            full_url = urljoin(listing_url, href)
            product_urls.append(full_url)

    # Handle pagination
    next_page = soup.select_one('a.next-page')
    if next_page:
        next_url = urljoin(listing_url, next_page['href'])
        product_urls.extend(get_listing_urls(next_url))

    return product_urls

Step 4: Clean the data

Raw extracted data is messy. Clean it before storage:

Problem Example Fix
Extra whitespace " Widget Pro 3000 " .strip(), collapse internal spaces
HTML entities "Widget & Gadget" html.unescape()
Currency symbols in numbers "$49.99" Remove $, convert to float
Inconsistent formats "49.99", "49,99", "$49.99" Normalise to one format
Missing values None, empty string, "N/A" Standardise to None or empty
Unicode issues Mixed encodings, mojibake Detect and convert to UTF-8

Step 5: Store the data

Format When to use
CSV Small to medium datasets, spreadsheet-compatible, human-readable
JSON Nested data, API consumption, flexible schema
SQLite Structured queries, deduplication, relationships
PostgreSQL/MySQL Production use, multiple users, large datasets
Parquet Large analytical datasets, column-oriented queries

Tools

Libraries

Tool Language Best for
Beautiful Soup Python HTML parsing, simple scraping
lxml Python Fast HTML/XML parsing, XPath
Scrapy Python Large-scale scraping with built-in pipeline
Playwright Python/JS JavaScript-rendered pages
Selenium Python/JS/Java Browser automation, dynamic pages
Cheerio JavaScript HTML parsing in Node.js
Puppeteer JavaScript Headless Chrome automation

Managed services

Service What it does
Apify Cloud scraping platform with pre-built actors
Bright Data Proxy network + scraping tools
ScrapingBee API for scraping with JavaScript rendering
Diffbot AI-powered structured data extraction
Import.io Visual web data extraction

No-code tools

Tool How it works
Octoparse Visual point-and-click scraper
Browse AI Train a robot by showing it what to extract
ParseHub Visual scraping with pagination handling
Web Scraper (Chrome extension) Build scrapers in the browser
  • Robots.txt. Check and respect the site's robots.txt file.
  • Terms of service. Many sites prohibit scraping in their ToS.
  • Rate limiting. Do not overload servers. Add delays between requests.
  • Copyright. Extracted data may be copyrighted. Extraction for analysis and personal use is generally more defensible than republication.
  • Personal data. Extracting and storing personal data must comply with GDPR, CCPA and other privacy regulations.

See Is Web Scraping Legal and Scraping Ethics Best Practices.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)