Proxy Rotation for Web Scraping: Types, Setup and Best Practices
By Email ExtractorPublished 8 min read
On this page
Why Proxies Matter for Web Scraping
When scraping websites, sending many requests from a single IP address triggers rate limits and blocks. Proxies route your requests through different IP addresses, making your traffic appear to come from multiple sources:
Problem without proxies
How proxies solve it
IP-based rate limiting
Requests come from different IPs; each stays within limits
IP blocking
Blocked IP is replaced by a fresh one
Geographic restrictions
Proxies in the target country access local content
Fingerprinting
Different IPs make requests harder to correlate
Proxy Types
By source
Type
How it works
Speed
Cost
Detection risk
Datacenter
IPs from cloud hosting providers (AWS, GCP, Hetzner, etc.)
Very fast
Low ($0.50-3 per IP/month)
High (easily identified as non-residential)
Residential
IPs from real ISP customers (obtained through SDK partnerships or peer networks)
Moderate
High ($5-15 per GB)
Low (appears as normal home user)
ISP/static residential
Datacenter-hosted IPs registered to an ISP
Fast
Medium ($2-5 per IP/month)
Medium (registered as residential but hosted in datacenter)
Mobile
IPs from mobile carriers (3G/4G/5G connections)
Slow
Very high ($15-30+ per GB)
Very low (hardest to detect; shared by many users)
By rotation model
Model
How it works
Best for
Rotating
New IP for every request or every N seconds
High-volume scraping; avoiding per-IP limits
Sticky sessions
Same IP maintained for a set duration (e.g., 10 minutes)
Sessions requiring login or multi-page navigation
Static
Fixed IP that does not change
Long-running tasks, whitelisted access
Backconnect
Single endpoint that rotates IPs on the backend
Simplest integration; provider handles rotation
By protocol
Protocol
Use case
Notes
HTTP/HTTPS
Standard web scraping
Most common; supports most scraping tools
SOCKS5
Any TCP traffic; more flexible than HTTP
Supports more protocols; useful for non-HTTP scraping
Browser proxy
Configured in headless browsers (Puppeteer, Playwright)
Used with browser-based scraping
Proxy Rotation Strategies
Basic rotation
Strategy
Implementation
When to use
Round-robin
Cycle through proxy list sequentially
Simple, predictable distribution
Random
Select a random proxy for each request
Easy to implement; avoids patterns
Weighted random
Assign weights based on proxy performance
Favour faster, more reliable proxies
Least-recently-used
Use the proxy that has been idle longest
Maximise cooldown time between uses
Advanced rotation
Strategy
Implementation
When to use
Per-domain rotation
Different proxy pool per target domain
Each domain sees different IP patterns
Geographic matching
Use proxies from the same country as the target
Access geo-restricted content; appear local
Retry with new proxy
On block/CAPTCHA, retry with a different proxy
Handle blocks without losing the request
Session-based
Maintain same IP for related requests, rotate between sessions
Navigate multi-page flows
Time-based
Rotate proxy after N seconds regardless of requests
Limit exposure time per IP
Health-based
Remove failing proxies from rotation; re-check periodically
Maintain pool quality
Python implementation: Basic proxy rotation
import requests
import random
import time
from itertools import cycle
class ProxyRotator:
"""Manage proxy rotation for web scraping."""
def __init__(self, proxies):
self.proxies = proxies
self.working_proxies = list(proxies)
self.failed_proxies = {}
self.cycle = cycle(self.working_proxies)
def get_next_proxy(self, strategy='round_robin'):
"""Return the next proxy based on the selected strategy."""
if strategy == 'round_robin':
return next(self.cycle)
elif strategy == 'random':
return random.choice(self.working_proxies)
else:
return next(self.cycle)
def mark_failed(self, proxy, cooldown=300):
"""Remove a proxy from rotation temporarily."""
if proxy in self.working_proxies:
self.working_proxies.remove(proxy)
self.failed_proxies[proxy] = time.time() + cooldown
# Rebuild cycle without the failed proxy
self.cycle = cycle(self.working_proxies)
print(f"Proxy removed: {proxy} "
f"(cooldown: {cooldown}s, "
f"remaining: {len(self.working_proxies)})")
def restore_proxies(self):
"""Restore proxies whose cooldown has expired."""
now = time.time()
restored = []
for proxy, expiry in list(self.failed_proxies.items()):
if now >= expiry:
self.working_proxies.append(proxy)
restored.append(proxy)
del self.failed_proxies[proxy]
if restored:
self.cycle = cycle(self.working_proxies)
print(f"Restored {len(restored)} proxies. "
f"Pool size: {len(self.working_proxies)}")
def fetch(self, url, max_retries=3, delay=1):
"""Fetch a URL using proxy rotation with retries."""
self.restore_proxies()
for attempt in range(max_retries):
if not self.working_proxies:
print("No working proxies available.")
return None
proxy = self.get_next_proxy()
proxy_dict = {
'http': f'http://{proxy}',
'https': f'http://{proxy}'
}
try:
response = requests.get(
url,
proxies=proxy_dict,
timeout=15,
headers={
'User-Agent': (
'Mozilla/5.0 (Windows NT 10.0; Win64; x64) '
'AppleWebKit/537.36')
}
)
if response.status_code == 200:
return response
elif response.status_code in (403, 429, 503):
self.mark_failed(proxy)
time.sleep(delay)
else:
return response
except requests.RequestException:
self.mark_failed(proxy)
time.sleep(delay)
return None
# Usage
proxies = [
'user:pass@proxy1.example.com:8080',
'user:pass@proxy2.example.com:8080',
'user:pass@proxy3.example.com:8080',
]
rotator = ProxyRotator(proxies)
response = rotator.fetch('https://example.com/page')
Proxy Provider Comparison
Major proxy providers
Provider
Proxy types
Residential pool size
Pricing model
Minimum plan
Bright Data (formerly Luminati)
Datacenter, residential, ISP, mobile
72M+ IPs
Per GB (residential) or per IP (datacenter)
Varies; pay-as-you-go available
Oxylabs
Datacenter, residential, ISP, mobile
100M+ IPs
Per GB (residential) or per IP (datacenter)
Varies; monthly plans
Smartproxy
Datacenter, residential, ISP, mobile
55M+ IPs
Per GB (residential) or per request
Lower entry point than some competitors
IPRoyal
Datacenter, residential, ISP, mobile
2M+ IPs
Per GB or per IP
Pay-as-you-go available
SOAX
Residential, mobile, ISP
155M+ IPs
Per GB with port-based pricing
Monthly plans
Webshare
Datacenter, residential
Smaller pool
Per proxy per month
Low-cost option
Cost comparison
Proxy type
Typical cost
Cost per 1M requests (estimated)
Datacenter (shared)
$0.50-2/IP/month
$5-20 (assuming you own or rent a small pool)
Datacenter (dedicated)
$2-5/IP/month
$20-50
Residential (pay-per-GB)
$5-15/GB
$50-150 (depends on page size)
ISP/static residential
$2-5/IP/month
$20-50
Mobile
$15-30/GB
$150-300
Choosing the right proxy type
Scenario
Recommended type
Rationale
Scraping public data at high volume
Datacenter
Low cost; fast speed; acceptable detection risk for public data
Scraping sites with aggressive anti-bot
Residential
Appears as real users; low detection
Scraping geo-restricted content
Residential (targeted location)
Access content as if from that country
Navigating login-protected sessions
ISP/static residential with sticky sessions
Maintains session; appears residential
Scraping social media platforms
Residential or mobile
Aggressive detection on these platforms
Scraping search engine results
Residential or specialised SERP API
Search engines actively block datacenter IPs
Budget-constrained scraping
Datacenter
Lowest cost; accept higher block rates
Anti-Detection Best Practices
Proxies alone do not prevent detection. Combine proxies with other techniques:
Technique
What it does
How to implement
User-Agent rotation
Different browser identities per request
Maintain a list of real user agents; rotate with proxy
Random delays of 2-10 seconds; avoid exact intervals
Cookie handling
Maintain cookies within sessions
Use session objects; clear cookies between sessions
TLS fingerprinting
Browser-like TLS handshake
Use browser-based tools or TLS fingerprint libraries
JavaScript rendering
Execute JavaScript like a real browser
Use headless browsers (Puppeteer, Playwright)
CAPTCHA handling
Solve CAPTCHAs when encountered
CAPTCHA solving services or manual solving
Common mistakes that reveal scrapers
Mistake
Why it gets detected
Fix
Sequential page access
Real users do not visit page 1, 2, 3, 4, 5 in exact order
Add random navigation patterns
Exact timing between requests
Real users have variable timing
Use random delays with normal distribution
No CSS/image requests
Real browsers load all page resources
Use a real browser or request associated resources
Missing cookies
Real browsers maintain cookies
Accept and send cookies
Same proxy for too many requests
Exceeds normal user behaviour from one IP
Rotate more frequently
Datacenter IP ranges
Identified by IP reputation databases
Use residential proxies for sensitive targets
Missing JavaScript execution
Page may require JS to render content
Use headless browser
Wrong geographic proxy
Proxy location does not match expected user location
Use geo-targeted proxies
Proxy Infrastructure Architecture
Small scale (under 10K requests/day)
Component
Implementation
Proxy source
10-50 datacenter proxies or a small residential plan
Rotation
Simple round-robin or random
Retry logic
Retry with a different proxy on failure
Monitoring
Log success/failure rates per proxy
Medium scale (10K-1M requests/day)
Component
Implementation
Proxy source
Mix of datacenter and residential; 100+ IPs
Rotation
Health-based rotation with cooldown
Retry logic
Multi-level retry with proxy type escalation (datacenter first, residential on block)
Queue
Request queue to manage throughput
Monitoring
Dashboard tracking success rates, latency, cost per request
IP management
Automatic proxy health checks; remove and restore
Large scale (1M+ requests/day)
Component
Implementation
Proxy source
Multiple providers; thousands of IPs; mix of types
Rotation
Sophisticated routing based on target domain, response analysis and historical performance
Retry logic
Automated retry with escalation, backoff and circuit breakers
Queue
Distributed queue (Redis, RabbitMQ) for request management
Monitoring
Real-time alerting on success rates, block rates, cost
IP management
Automated pool management with health scores
Cost optimisation
Route requests to cheapest working proxy type
Failover
Multiple proxy providers with automatic failover
Legal and Ethical Considerations
Consideration
Guidance
Terms of service
Review target site's ToS; some explicitly prohibit proxy use
Robots.txt
Respect crawl directives even when using proxies
Rate limiting
Proxies do not exempt you from being a polite scraper; maintain reasonable request rates
Data protection
GDPR and other laws apply to data collected through proxies
Residential proxy ethics
Understand how residential IPs are sourced; some providers use peer networks with varying levels of user consent
Transparency
If asked by a website operator, be honest about your scraping activities
Proportionality
Collect only the data you need; do not scrape entire sites unnecessarily
Processing Scraped Data
After collecting data through proxied scraping, the raw HTML or downloaded files often contain email addresses and contact information mixed with other content. Upload scraped files to Email Extractor to extract email addresses from HTML, JSON, XML, CSV or text files. The tool handles multiple file formats, so you can process scraped output regardless of the format your scraping pipeline saves it in.