Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Is Web Scraping Legal? What You Need to Know in 2026

On this page

The Short Answer

Web scraping is not inherently illegal, but it is not automatically legal either. The legality depends on what you scrape, how you scrape it, where the data comes from, and what you do with it afterwards. The legal landscape varies by country and continues to evolve.

This article provides a general overview. It is not legal advice. Consult a lawyer for guidance on your specific situation.

Computer Fraud and Abuse Act (CFAA) -- United States

The CFAA is a federal law that prohibits unauthorized access to computer systems. The central question for scraping is whether accessing a publicly available website counts as "unauthorized access."

The hiQ Labs v. LinkedIn case (2022): The US Ninth Circuit Court of Appeals ruled that scraping publicly available data from LinkedIn did not violate the CFAA. The court reasoned that data visible to any member of the public is not behind an authorization barrier, so accessing it is not "without authorization" under the CFAA.

What this means in practice:

  • Scraping public data from public websites is less likely to violate the CFAA.
  • Scraping data behind a login, paywall or other access control is riskier and more likely to be considered unauthorized access.
  • Circumventing technical barriers (IP blocks, CAPTCHAs, rate limits) after being told to stop could be treated as unauthorized access.

General Data Protection Regulation (GDPR) -- EU/EEA

GDPR regulates the processing of personal data of individuals in the EU/EEA, regardless of where the scraper is located.

Key requirements:

  • Lawful basis: You need a lawful basis to collect personal data. For scraped data, "legitimate interest" is the most commonly cited basis, but it requires a balancing test against the data subject's rights.
  • Transparency: You must inform data subjects that you have collected their data and explain how you will use it.
  • Purpose limitation: Data collected for one purpose cannot be used for an unrelated purpose.
  • Data minimisation: Only collect the data you actually need.
  • Right to erasure: Individuals can request that their data be deleted.

Scraping email addresses of EU residents and using them for unsolicited outreach without meeting these requirements is a GDPR violation regardless of whether the website is public.

For more on GDPR, see GDPR and Email Extraction.

Terms of Service

Most websites include terms of service (ToS) that explicitly prohibit scraping, automated data collection or commercial use of the site's content.

Legal weight of ToS:

  • Violating a website's ToS is generally a breach of contract, not a criminal offence.
  • Courts have varied on whether ToS are enforceable against scrapers who never explicitly agreed to them (browsewrap agreements).
  • The hiQ v. LinkedIn case suggested that ToS alone cannot prevent scraping of public data, but this does not mean ToS violations are consequence-free. The site operator can still pursue civil claims.

Copyright

Website content may be protected by copyright. Scraping and republishing copyrighted content (articles, images, databases) without permission can constitute copyright infringement.

Facts vs. creative works:

  • Individual facts (a business name, an email address, a phone number) are generally not copyrightable.
  • The creative arrangement and selection of facts in a database may be protected (in the EU, the Database Directive provides sui generis database rights).
  • Scraping factual data and reorganising it independently is less risky than copying a database wholesale.

CAN-SPAM Act (United States)

CAN-SPAM does not prohibit scraping email addresses. It regulates what you do after you have them. If you send commercial emails to scraped addresses, you must include a physical postal address, a clear unsubscribe mechanism and honest subject lines. CAN-SPAM does not require prior opt-in for commercial email in the US, but non-compliance carries penalties.

See CAN-SPAM and Email Extraction.

CASL (Canada)

Canada's Anti-Spam Legislation is stricter than CAN-SPAM. CASL requires express or implied consent before sending commercial electronic messages. Scraping an email address does not create implied consent. Sending to scraped Canadian addresses without consent violates CASL.

See CASL and Email Extraction.

Public vs. private data

Scraping data that anyone can see without logging in is generally lower risk. Scraping data behind a login, behind a paywall, in a private group or accessible only after agreeing to terms is higher risk.

Respecting robots.txt

robots.txt is a file that tells automated systems which parts of a website the operator prefers not to be crawled. While robots.txt is not legally binding, ignoring it can be used as evidence that you knew the site operator did not want their content scraped. Respecting robots.txt demonstrates good faith.

Rate limiting and server impact

Scraping aggressively enough to slow down or crash a website can constitute a denial-of-service attack, which is illegal in most jurisdictions. Sending requests at a reasonable rate, respecting rate limits and backing off when a server returns error codes are both good practice and risk mitigation.

What you do with the data

Scraping data for personal research is very different from scraping data to build a competing product, resell the data commercially or send mass unsolicited emails. The intended use is a major factor in legal risk.

Cease and desist notices

If a website operator sends you a cease and desist letter asking you to stop scraping, continuing to scrape after receiving it significantly increases your legal risk. In the CFAA context, continuing after a clear objection from the site operator could be treated as exceeding authorized access.

Practical Guidelines

Lower risk

  • Scraping publicly available data from public websites.
  • Respecting robots.txt and rate limits.
  • Collecting factual data (names, titles, email addresses on public pages).
  • Using the data for personal research or internal analysis.
  • Honouring opt-out and deletion requests.

Higher risk

  • Scraping data behind logins or paywalls.
  • Ignoring robots.txt and cease-and-desist notices.
  • Overwhelming servers with aggressive request rates.
  • Scraping and republishing copyrighted content.
  • Collecting personal data of EU residents without a lawful basis.
  • Using scraped emails for mass unsolicited outreach without compliance.

How Email Extractor Differs from Web Scraping

Email Extractor is a client-side tool that extracts email addresses from files you already have. It processes text files, CSVs, PDFs, spreadsheets, HTML files and other documents locally in your browser. It does not visit websites, bypass access controls or interact with servers on your behalf.

The webpage extraction feature sends URLs to the server for processing, but this is a separate workflow from file-based extraction. It retrieves the publicly visible content of the pages you provide.

The legal considerations around how you obtained the source files and how you use the extracted emails still apply. Email Extractor does not change your legal obligations regarding the data.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)