Web Scraping Ethics and Best Practices: A Responsible Guide
On this page
Why Ethics Matter in Web Scraping
Web scraping is a powerful tool for gathering business data, monitoring competitors, conducting research and building contact lists. But the ability to extract data from websites comes with responsibilities.
Irresponsible scraping can overwhelm servers, violate privacy, break laws and get your IP addresses and domains blocked. Responsible scraping protects both you and the websites you interact with.
This guide covers the ethical principles and practical best practices for scraping data from the web.
Respecting robots.txt
What it is
The robots.txt file sits at the root of every website (e.g., https://example.com/robots.txt). It tells automated agents (bots, crawlers, scrapers) which parts of the site they may and may not access.
How to read it
User-agent: *
Disallow: /private/
Disallow: /api/
Crawl-delay: 10
User-agent: Googlebot
Allow: /
This file says:
- All bots should avoid /private/ and /api/ directories.
- All bots should wait 10 seconds between requests.
- Googlebot has full access (Google's crawler gets special treatment from this site).
Should you obey it?
robots.txt is a convention, not a technical barrier. Nothing prevents your scraper from accessing disallowed paths.
However:
- Ignoring robots.txt can violate the website's terms of service.
- In some jurisdictions, ignoring robots.txt has been cited as evidence of unauthorised access.
- Respecting it demonstrates good faith, which matters if a site owner ever questions your activity.
Best practice: Always check and respect robots.txt. If the data you need is in a disallowed section, find another source or contact the site owner for permission.
Rate Limiting
Why it matters
Every request your scraper makes consumes server resources. A scraper that sends 100 requests per second can degrade performance for real users or even take a small site offline.
This is not theoretical. Small business websites, personal blogs and community forums run on limited hosting. A sudden burst of scraping traffic can cost the site owner money and disrupt their visitors.
How to rate limit
Add delays between requests. A delay of 1-5 seconds between requests is reasonable for most sites. For small sites, 5-10 seconds is more appropriate.
Respect Crawl-delay. If robots.txt specifies a Crawl-delay, honour it.
Back off on errors. If you receive 429 (Too Many Requests) or 503 (Service Unavailable) responses, stop immediately and wait before retrying.
Scrape during off-peak hours. If the site has heavy traffic during business hours, scrape during evenings or weekends.
Limit concurrent connections. Do not open 50 simultaneous connections to the same server. Use 1-3 concurrent connections maximum.
Example rate limiting
import time
import random
# Between requests
time.sleep(random.uniform(2, 5)) # Random 2-5 second delay
Random delays are better than fixed delays because they more closely resemble human browsing behaviour and are less likely to trigger anti-bot systems.
Terms of Service
What to check
Most websites have Terms of Service (ToS) or Terms of Use that describe what you may and may not do with the site and its content.
Common restrictions:
- No automated access or data collection.
- No reproduction or redistribution of content.
- No commercial use of scraped data.
- No accessing the site in a way that could damage, disable or impair it.
The legal landscape
The legality of scraping depends on:
- What you scrape (public data vs protected data).
- How you access it (compliance with robots.txt and ToS).
- What you do with the data (personal use vs commercial redistribution).
- Your jurisdiction and the site's jurisdiction.
Courts in different countries have reached different conclusions. Some key cases (up to 2025):
- US courts have generally held that scraping publicly available data is not a violation of the Computer Fraud and Abuse Act (hiQ Labs v. LinkedIn).
- EU courts have applied GDPR to scraped personal data regardless of whether the data was publicly available.
- Some courts have enforced ToS prohibitions on scraping through breach of contract claims.
For a detailed analysis, see Is Web Scraping Legal?.
Best practice: Read the ToS before scraping. If the ToS explicitly prohibits automated access, consider whether you have a legitimate basis to proceed, or look for alternative data sources.
Data Protection
Personal data
Email addresses, names, phone numbers and other identifiers are personal data under GDPR, CCPA and similar laws. Scraping personal data from websites does not exempt you from data protection obligations.
Key obligations:
- Purpose limitation. Use scraped personal data only for the specific purpose you collected it for.
- Data minimisation. Only scrape the data you actually need. Do not collect entire profiles when you only need email addresses.
- Storage limitation. Do not store personal data indefinitely. Delete it when you no longer need it.
- Security. Protect scraped data with appropriate security measures.
- Rights compliance. Be prepared to respond to access, deletion and correction requests from data subjects.
Consent and legal basis
GDPR requires a legal basis for processing personal data. For scraped data, the most commonly cited bases are:
- Legitimate interest. Your business interest in the data outweighs the individual's privacy interest. Requires a legitimate interest assessment.
- Consent. You have the individual's consent. This is difficult to obtain for scraped data by definition.
See GDPR Guide and CCPA Guide.
What to avoid
- Scraping data behind login walls without authorisation.
- Scraping sensitive personal data (health, political opinions, religious beliefs).
- Scraping data about children.
- Combining scraped data with other sources to build detailed profiles without a clear legal basis.
- Selling scraped personal data without appropriate disclosures and opt-out mechanisms.
Server Impact
Monitor your impact
Good scrapers monitor the effect of their activity on the target server:
- Response times. If response times increase during your scraping, slow down.
- Error rates. A spike in 500 errors means you are stressing the server.
- Status codes. 429 responses mean you are being rate-limited.
Caching
Cache pages you have already fetched. If your scraper revisits the same URLs during a session, serve the cached version instead of hitting the server again.
Use APIs when available
Many websites offer APIs for programmatic data access. APIs are designed for automated use and are more efficient than scraping HTML.
Check for APIs before scraping. The data you need may be available through a public API with clear usage limits and documentation.
Identify your scraper
Set a descriptive User-Agent string that identifies your scraper and includes contact information:
User-Agent: CompanyNameBot/1.0 (contact@example.com)
This lets site operators contact you if your scraper causes problems, rather than blocking you outright.
Data Quality and Integrity
Verify accuracy
Scraped data contains errors. Websites have typos, outdated information and inconsistent formatting. Always validate and clean scraped data before using it.
For scraped email addresses, run them through Email Extractor to normalise formatting and remove duplicates, then verify with an email verification service.
Do not misrepresent
If you scrape reviews, prices, product information or other content, do not alter it, present it as your own, or take it out of context in misleading ways.
Attribution
If you publish or share scraped content, attribute the source. This is both ethical and may be legally required depending on the content's copyright status.
Checklist for Ethical Scraping
Before starting a scraping project, run through this checklist:
- Check robots.txt. Does it allow access to the pages you want to scrape?
- Read the Terms of Service. Does the site prohibit automated data collection?
- Identify an API. Is there an official API for the data you need?
- Assess personal data. Will you be scraping personal information? If so, do you have a legal basis?
- Set rate limits. Have you configured delays between requests?
- Set up monitoring. Can you detect if your scraper is impacting the server?
- Configure caching. Will your scraper avoid re-fetching pages it has already retrieved?
- Set a User-Agent. Does your scraper identify itself and provide contact information?
- Plan data handling. How will you store, protect and eventually delete the scraped data?
- Document your decisions. Keep a record of what you scraped, when, and your reasoning for compliance purposes.
When Not to Scrape
Sometimes scraping is not the right approach:
- An API exists. Use it instead.
- The data is behind a login. Accessing data behind authentication without permission is likely unauthorised access.
- The site explicitly prohibits it and you have no legal basis to override that prohibition.
- Your scraping would harm the site's performance. Find a less impactful method or a different source.
- You need the data for purposes the data subjects would not expect. Respect privacy.