Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Rate Limiting and Politeness: How to Scrape Without Burning Bridges

On this page

Why Politeness Matters

Aggressive scraping harms both the target website and the scraper:

For the target website:

  • Server overload. Too many requests per second can slow the site for real users or trigger auto-scaling costs for the site owner.
  • Bandwidth consumption. Scraping every page on a site consumes bandwidth that the site pays for.
  • Security alerts. Aggressive scraping patterns trigger security systems, creating work for the site's operations team.

For the scraper:

  • IP blocking. The most common consequence of aggressive scraping is getting your IP address blocked.
  • Legal risk. Overwhelming a server or violating terms of service strengthens a legal case against you.
  • Relationship damage. If the target is a potential customer, partner or industry peer, aggressive scraping damages the relationship.
  • Data quality. Blocked or throttled requests return errors instead of data, reducing the completeness of your results.

Polite scraping is not just ethical. It is practical. A scraper that runs slowly and consistently over days collects more complete data than one that runs fast and gets blocked after 200 pages.

Robots.txt

What it is

robots.txt is a file at the root of a website (e.g., https://example.com/robots.txt) that tells crawlers which parts of the site they may and may not access. It is a voluntary protocol: there is no technical enforcement. But respecting it is the foundation of polite scraping.

Reading robots.txt

User-agent: *
Disallow: /admin/
Disallow: /private/
Disallow: /api/
Crawl-delay: 10

User-agent: Googlebot
Allow: /
Crawl-delay: 1

Sitemap: https://example.com/sitemap.xml
Directive Meaning
User-agent: * Rules for all crawlers
Disallow: /admin/ Do not access any URL starting with /admin/
Allow: / Override a broader Disallow
Crawl-delay: 10 Wait at least 10 seconds between requests
Sitemap: Location of the XML sitemap

Implementing robots.txt compliance in Python

from urllib.robotparser import RobotFileParser
from urllib.parse import urlparse

class PoliteRobotChecker:
    def __init__(self):
        self.parsers = {}

    def can_fetch(self, url, user_agent="*"):
        """Check if a URL is allowed by robots.txt."""
        parsed = urlparse(url)
        domain = parsed.netloc

        if domain not in self.parsers:
            robots_url = f"{parsed.scheme}://{domain}/robots.txt"
            rp = RobotFileParser()
            rp.set_url(robots_url)
            try:
                rp.read()
            except Exception:
                # If robots.txt is unreachable, assume access is allowed
                rp = RobotFileParser()
            self.parsers[domain] = rp

        return self.parsers[domain].can_fetch(user_agent, url)

    def crawl_delay(self, domain, user_agent="*"):
        """Get the crawl delay for a domain."""
        if domain in self.parsers:
            delay = self.parsers[domain].crawl_delay(user_agent)
            return delay if delay else None
        return None

When robots.txt says "no"

If robots.txt disallows the pages you want to scrape:

  1. Respect it. This is the polite and legally safer choice.
  2. Look for alternatives. The site may have a public API, data feed or export option.
  3. Ask for permission. Contact the site owner and explain what data you need and why.
  4. Use manual collection. Visit the pages manually and copy the content you need. For email extraction, paste the copied text into Email Extractor.

Rate Limiting Strategies

Fixed delay

The simplest approach: wait a fixed number of seconds between every request.

import time

def scrape_with_fixed_delay(urls, delay=2):
    for url in urls:
        response = fetch(url)
        process(response)
        time.sleep(delay)
Delay Requests per minute Use case
1 second 60 Light scraping of large, high-traffic sites
2 seconds 30 General-purpose, good default
5 seconds 12 Sites with limited capacity or strict rate limits
10 seconds 6 When robots.txt specifies crawl-delay: 10
30 seconds 2 Very conservative, when you want to be invisible

Random delay

Adding randomness makes the request pattern look less automated:

import random
import time

def scrape_with_random_delay(urls, min_delay=1, max_delay=3):
    for url in urls:
        response = fetch(url)
        process(response)
        time.sleep(random.uniform(min_delay, max_delay))

Per-domain throttling

When scraping across many different websites, throttle per domain rather than globally. This lets you maintain high overall throughput while being polite to each individual site:

import time
from collections import defaultdict
from urllib.parse import urlparse

class DomainThrottler:
    def __init__(self, default_delay=2):
        self.last_request = defaultdict(float)
        self.default_delay = default_delay

    def wait(self, url):
        domain = urlparse(url).netloc
        elapsed = time.time() - self.last_request[domain]
        if elapsed < self.default_delay:
            time.sleep(self.default_delay - elapsed)
        self.last_request[domain] = time.time()

Adaptive throttling

Adjust speed based on the server's response:

class AdaptiveThrottler:
    def __init__(self, initial_delay=1, max_delay=30):
        self.delays = defaultdict(lambda: initial_delay)
        self.max_delay = max_delay

    def on_success(self, domain):
        # Slightly decrease delay on success (minimum 1 second)
        self.delays[domain] = max(1, self.delays[domain] * 0.9)

    def on_rate_limit(self, domain):
        # Double delay on rate limit
        self.delays[domain] = min(
            self.max_delay, self.delays[domain] * 2
        )

    def on_error(self, domain):
        # Increase delay on error
        self.delays[domain] = min(
            self.max_delay, self.delays[domain] * 1.5
        )

Concurrent scraping with per-domain limits

For maximum throughput with politeness, process multiple domains in parallel while limiting to one request per domain at a time:

import asyncio
import aiohttp
from collections import defaultdict

class PoliteCrawler:
    def __init__(self, delay=2, max_concurrent=10):
        self.delay = delay
        self.semaphore = asyncio.Semaphore(max_concurrent)
        self.domain_locks = defaultdict(asyncio.Lock)

    async def fetch(self, session, url):
        domain = urlparse(url).netloc
        async with self.semaphore:
            async with self.domain_locks[domain]:
                async with session.get(url) as response:
                    content = await response.text()
                await asyncio.sleep(self.delay)
                return content

This pattern lets you scrape 10 domains simultaneously, each at a polite 2-second interval, for an effective throughput of 5 pages per second across all domains combined.

Crawl Budget Management

A crawl budget is the total number of pages you plan to scrape from a site. Setting and enforcing crawl budgets prevents runaway crawls.

Setting budgets

Site type Suggested budget Why
Small business site (under 100 pages) 10-20 pages Contact + team + about is usually enough
Medium business site (100-1,000 pages) 50-100 pages Focus on directories and contact sections
Large corporate site (1,000+ pages) 100-200 pages Target specific sections, not the whole site
Directory or listing site 500-1,000 pages Paginated listings require more pages

Prioritising pages

Not all pages are equally valuable. Prioritise:

  1. Contact and team pages (highest probability of email addresses).
  2. Staff directories and department pages.
  3. Press and media pages.
  4. About and leadership pages.
  5. Blog author pages.

Skip:

  • Blog post content (rarely contains contact emails).
  • Product pages (no contact information).
  • Legal pages (terms, privacy policy).
  • Image and asset URLs.

Error Handling

HTTP response codes

Code Meaning Action
200 Success Process the page
301/302 Redirect Follow the redirect (most libraries do this automatically)
403 Forbidden You are blocked; stop scraping this domain
404 Not found Skip this URL
429 Too many requests Back off; increase delay significantly
500 Server error Retry once after a delay; skip if it persists
503 Service unavailable Server is overloaded or in maintenance; back off and retry later

Retry logic

import time
import requests

def fetch_with_retry(url, max_retries=3, initial_delay=5):
    """Fetch URL with exponential backoff retry."""
    delay = initial_delay
    for attempt in range(max_retries):
        try:
            response = requests.get(url, timeout=10)
            if response.status_code == 429:
                # Rate limited: back off
                retry_after = int(
                    response.headers.get("Retry-After", delay)
                )
                time.sleep(retry_after)
                delay *= 2
                continue
            if response.status_code == 503:
                # Server overloaded: back off
                time.sleep(delay)
                delay *= 2
                continue
            return response
        except requests.RequestException:
            if attempt < max_retries - 1:
                time.sleep(delay)
                delay *= 2
    return None

When to stop

Stop scraping a domain when:

Signal Action
Three consecutive 403 responses IP is blocked; stop
429 response with Retry-After > 60 seconds Come back later or use a different IP
Connection timeouts Server may be struggling; stop and try later
CAPTCHA challenge You have been identified as a bot; stop
robots.txt disallows your target paths Respect the restriction

Identifying Your Scraper

Some scrapers identify themselves in the User-Agent string:

User-Agent: CompanyBot/1.0 (https://example.com/bot; bot@example.com)

Advantages:

  • Transparent. Site owners can contact you if they have concerns.
  • Can be specifically allowed in robots.txt.
  • Builds trust for ongoing data relationships.

Disadvantages:

  • Easy to block by User-Agent string.
  • Not suitable for all scraping use cases.

For business-to-business contact data collection, transparency is generally the better approach. You are collecting public business information for a legitimate purpose, and hiding your identity creates more risk than it avoids.

Alternatives to Scraping

When polite scraping is still too burdensome on target sites, consider alternatives:

Alternative When to use
Public APIs When the site offers structured data access
Data providers When someone already aggregates the data you need
RSS feeds For monitoring new content
Google Cache / search results For accessing page content without hitting the site directly
Manual collection + Email Extractor For small-scale collection from high-value targets

For extracting emails from files you already have (PDFs, spreadsheets, exported CSVs, HTML files), use Email Extractor directly. No scraping needed: upload the files, extract and deduplicate in your browser.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)