Rate Limiting and Politeness: How to Scrape Without Burning Bridges
On this page
Why Politeness Matters
Aggressive scraping harms both the target website and the scraper:
For the target website:
- Server overload. Too many requests per second can slow the site for real users or trigger auto-scaling costs for the site owner.
- Bandwidth consumption. Scraping every page on a site consumes bandwidth that the site pays for.
- Security alerts. Aggressive scraping patterns trigger security systems, creating work for the site's operations team.
For the scraper:
- IP blocking. The most common consequence of aggressive scraping is getting your IP address blocked.
- Legal risk. Overwhelming a server or violating terms of service strengthens a legal case against you.
- Relationship damage. If the target is a potential customer, partner or industry peer, aggressive scraping damages the relationship.
- Data quality. Blocked or throttled requests return errors instead of data, reducing the completeness of your results.
Polite scraping is not just ethical. It is practical. A scraper that runs slowly and consistently over days collects more complete data than one that runs fast and gets blocked after 200 pages.
Robots.txt
What it is
robots.txt is a file at the root of a website (e.g., https://example.com/robots.txt) that tells crawlers which parts of the site they may and may not access. It is a voluntary protocol: there is no technical enforcement. But respecting it is the foundation of polite scraping.
Reading robots.txt
User-agent: *
Disallow: /admin/
Disallow: /private/
Disallow: /api/
Crawl-delay: 10
User-agent: Googlebot
Allow: /
Crawl-delay: 1
Sitemap: https://example.com/sitemap.xml
| Directive | Meaning |
|---|---|
| User-agent: * | Rules for all crawlers |
| Disallow: /admin/ | Do not access any URL starting with /admin/ |
| Allow: / | Override a broader Disallow |
| Crawl-delay: 10 | Wait at least 10 seconds between requests |
| Sitemap: | Location of the XML sitemap |
Implementing robots.txt compliance in Python
from urllib.robotparser import RobotFileParser
from urllib.parse import urlparse
class PoliteRobotChecker:
def __init__(self):
self.parsers = {}
def can_fetch(self, url, user_agent="*"):
"""Check if a URL is allowed by robots.txt."""
parsed = urlparse(url)
domain = parsed.netloc
if domain not in self.parsers:
robots_url = f"{parsed.scheme}://{domain}/robots.txt"
rp = RobotFileParser()
rp.set_url(robots_url)
try:
rp.read()
except Exception:
# If robots.txt is unreachable, assume access is allowed
rp = RobotFileParser()
self.parsers[domain] = rp
return self.parsers[domain].can_fetch(user_agent, url)
def crawl_delay(self, domain, user_agent="*"):
"""Get the crawl delay for a domain."""
if domain in self.parsers:
delay = self.parsers[domain].crawl_delay(user_agent)
return delay if delay else None
return None
When robots.txt says "no"
If robots.txt disallows the pages you want to scrape:
- Respect it. This is the polite and legally safer choice.
- Look for alternatives. The site may have a public API, data feed or export option.
- Ask for permission. Contact the site owner and explain what data you need and why.
- Use manual collection. Visit the pages manually and copy the content you need. For email extraction, paste the copied text into Email Extractor.
Rate Limiting Strategies
Fixed delay
The simplest approach: wait a fixed number of seconds between every request.
import time
def scrape_with_fixed_delay(urls, delay=2):
for url in urls:
response = fetch(url)
process(response)
time.sleep(delay)
| Delay | Requests per minute | Use case |
|---|---|---|
| 1 second | 60 | Light scraping of large, high-traffic sites |
| 2 seconds | 30 | General-purpose, good default |
| 5 seconds | 12 | Sites with limited capacity or strict rate limits |
| 10 seconds | 6 | When robots.txt specifies crawl-delay: 10 |
| 30 seconds | 2 | Very conservative, when you want to be invisible |
Random delay
Adding randomness makes the request pattern look less automated:
import random
import time
def scrape_with_random_delay(urls, min_delay=1, max_delay=3):
for url in urls:
response = fetch(url)
process(response)
time.sleep(random.uniform(min_delay, max_delay))
Per-domain throttling
When scraping across many different websites, throttle per domain rather than globally. This lets you maintain high overall throughput while being polite to each individual site:
import time
from collections import defaultdict
from urllib.parse import urlparse
class DomainThrottler:
def __init__(self, default_delay=2):
self.last_request = defaultdict(float)
self.default_delay = default_delay
def wait(self, url):
domain = urlparse(url).netloc
elapsed = time.time() - self.last_request[domain]
if elapsed < self.default_delay:
time.sleep(self.default_delay - elapsed)
self.last_request[domain] = time.time()
Adaptive throttling
Adjust speed based on the server's response:
class AdaptiveThrottler:
def __init__(self, initial_delay=1, max_delay=30):
self.delays = defaultdict(lambda: initial_delay)
self.max_delay = max_delay
def on_success(self, domain):
# Slightly decrease delay on success (minimum 1 second)
self.delays[domain] = max(1, self.delays[domain] * 0.9)
def on_rate_limit(self, domain):
# Double delay on rate limit
self.delays[domain] = min(
self.max_delay, self.delays[domain] * 2
)
def on_error(self, domain):
# Increase delay on error
self.delays[domain] = min(
self.max_delay, self.delays[domain] * 1.5
)
Concurrent scraping with per-domain limits
For maximum throughput with politeness, process multiple domains in parallel while limiting to one request per domain at a time:
import asyncio
import aiohttp
from collections import defaultdict
class PoliteCrawler:
def __init__(self, delay=2, max_concurrent=10):
self.delay = delay
self.semaphore = asyncio.Semaphore(max_concurrent)
self.domain_locks = defaultdict(asyncio.Lock)
async def fetch(self, session, url):
domain = urlparse(url).netloc
async with self.semaphore:
async with self.domain_locks[domain]:
async with session.get(url) as response:
content = await response.text()
await asyncio.sleep(self.delay)
return content
This pattern lets you scrape 10 domains simultaneously, each at a polite 2-second interval, for an effective throughput of 5 pages per second across all domains combined.
Crawl Budget Management
A crawl budget is the total number of pages you plan to scrape from a site. Setting and enforcing crawl budgets prevents runaway crawls.
Setting budgets
| Site type | Suggested budget | Why |
|---|---|---|
| Small business site (under 100 pages) | 10-20 pages | Contact + team + about is usually enough |
| Medium business site (100-1,000 pages) | 50-100 pages | Focus on directories and contact sections |
| Large corporate site (1,000+ pages) | 100-200 pages | Target specific sections, not the whole site |
| Directory or listing site | 500-1,000 pages | Paginated listings require more pages |
Prioritising pages
Not all pages are equally valuable. Prioritise:
- Contact and team pages (highest probability of email addresses).
- Staff directories and department pages.
- Press and media pages.
- About and leadership pages.
- Blog author pages.
Skip:
- Blog post content (rarely contains contact emails).
- Product pages (no contact information).
- Legal pages (terms, privacy policy).
- Image and asset URLs.
Error Handling
HTTP response codes
| Code | Meaning | Action |
|---|---|---|
| 200 | Success | Process the page |
| 301/302 | Redirect | Follow the redirect (most libraries do this automatically) |
| 403 | Forbidden | You are blocked; stop scraping this domain |
| 404 | Not found | Skip this URL |
| 429 | Too many requests | Back off; increase delay significantly |
| 500 | Server error | Retry once after a delay; skip if it persists |
| 503 | Service unavailable | Server is overloaded or in maintenance; back off and retry later |
Retry logic
import time
import requests
def fetch_with_retry(url, max_retries=3, initial_delay=5):
"""Fetch URL with exponential backoff retry."""
delay = initial_delay
for attempt in range(max_retries):
try:
response = requests.get(url, timeout=10)
if response.status_code == 429:
# Rate limited: back off
retry_after = int(
response.headers.get("Retry-After", delay)
)
time.sleep(retry_after)
delay *= 2
continue
if response.status_code == 503:
# Server overloaded: back off
time.sleep(delay)
delay *= 2
continue
return response
except requests.RequestException:
if attempt < max_retries - 1:
time.sleep(delay)
delay *= 2
return None
When to stop
Stop scraping a domain when:
| Signal | Action |
|---|---|
| Three consecutive 403 responses | IP is blocked; stop |
| 429 response with Retry-After > 60 seconds | Come back later or use a different IP |
| Connection timeouts | Server may be struggling; stop and try later |
| CAPTCHA challenge | You have been identified as a bot; stop |
| robots.txt disallows your target paths | Respect the restriction |
Identifying Your Scraper
Some scrapers identify themselves in the User-Agent string:
User-Agent: CompanyBot/1.0 (https://example.com/bot; bot@example.com)
Advantages:
- Transparent. Site owners can contact you if they have concerns.
- Can be specifically allowed in robots.txt.
- Builds trust for ongoing data relationships.
Disadvantages:
- Easy to block by User-Agent string.
- Not suitable for all scraping use cases.
For business-to-business contact data collection, transparency is generally the better approach. You are collecting public business information for a legitimate purpose, and hiding your identity creates more risk than it avoids.
Alternatives to Scraping
When polite scraping is still too burdensome on target sites, consider alternatives:
| Alternative | When to use |
|---|---|
| Public APIs | When the site offers structured data access |
| Data providers | When someone already aggregates the data you need |
| RSS feeds | For monitoring new content |
| Google Cache / search results | For accessing page content without hitting the site directly |
| Manual collection + Email Extractor | For small-scale collection from high-value targets |
For extracting emails from files you already have (PDFs, spreadsheets, exported CSVs, HTML files), use Email Extractor directly. No scraping needed: upload the files, extract and deduplicate in your browser.