Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

How to Extract Email Addresses from Saved Web Pages and HTML Files

On this page

When Saved Web Pages Are Useful

You can save any web page as an HTML file and upload it to Email Extractor. This works well for pages that contain email addresses in their source code, such as staff directories, contact pages, about pages and event listings.

Email Extractor treats .html and .htm files identically. Both extensions are supported.

The browser extension is often a faster alternative for live web pages, since it scans pages directly without the save step. Saved HTML files are useful when you want to process pages offline, archive them for later extraction, or work with pages that have already been saved.

Extraction Steps

Saving a web page as HTML

  1. Open the web page in your browser.
  2. Press Ctrl+S (Windows/Linux) or Cmd+S (macOS).
  3. Choose "Webpage, HTML Only" or "Web Page, Complete" depending on your browser. The "HTML Only" option is sufficient for email extraction.
  4. Save the file to your computer.

Extracting from the saved file

  1. Open Email Extractor.
  2. Select Text and files.
  3. Upload the saved .html or .htm file, or drag it into the upload area. Maximum 25 MB per file, 100 MB per batch.
  4. Click Extract emails.
  5. Review the results. Duplicates are removed using case-insensitive matching.
  6. Copy individual addresses or download as TXT, CSV or CSV with sources.

Example

A company team page saved as team.html contains:

<div class="team-member">
  <h3>Yuki Tanaka</h3>
  <p>Lead Engineer</p>
  <a href="mailto:yuki@example.com">yuki@example.com</a>
</div>
<div class="team-member">
  <h3>Priya Shah</h3>
  <p>Designer</p>
  <a href="mailto:priya@example.org">Contact Priya</a>
</div>
<!-- Internal: admin@example.net handles site queries -->

After extraction, Email Extractor returns three unique addresses:

Email Extractor scans the raw HTML source, so it finds addresses in mailto: links, HTML comments, meta tags and other markup that may not be visible on the rendered page.

What Email Extractor Reads

Page source text. All text content in the HTML, including visible paragraphs, headings, list items and table cells.

Mailto links. The href attribute of <a> tags that use mailto: URLs. The email address is extracted from the link target.

HTML attributes and comments. Addresses in attributes (such as data-email), HTML comments and meta tags are found during the source scan.

Multiple HTML files. You can upload several saved pages in one batch and extract from all of them together.

Limitations

No JavaScript execution. Email Extractor does not execute JavaScript in uploaded HTML files. If a web page loads email addresses dynamically using JavaScript (for example, assembling addresses from parts to prevent scraping, or loading them via an API call after the page renders), those addresses will not be present in the saved HTML source.

To check: open the saved HTML file directly in your browser (double-click it). If email addresses are missing from this local view, they were loaded by JavaScript and are not in the file. The browser extension may work better for these pages, since it scans the rendered page after JavaScript has executed.

No linked assets fetched. When you upload an HTML file, Email Extractor reads only that file. It does not fetch stylesheets, images, scripts or other files referenced by the HTML. This does not affect email extraction in most cases, since addresses are in the HTML text rather than in linked assets.

Snapshot, not live. A saved HTML file is a snapshot from the moment you saved it. If the page has been updated since, the saved file will not reflect those changes.

Obfuscated addresses. Some websites deliberately obfuscate email addresses to prevent automated collection. Common techniques include encoding the @ symbol as [at], splitting the address across multiple HTML elements, or rendering addresses as images. These will not be extracted from the saved HTML. If you can see an address on the rendered page but it does not appear in the source, it may be obfuscated or JavaScript-generated.

Saving Tips for Better Results

Use "HTML Only" when possible. The "Complete" save option downloads images and other assets alongside the HTML file. These extra files are not needed for email extraction and increase the file size unnecessarily.

Save after the page fully loads. Wait for the page to finish loading before saving. Some pages load content progressively, and saving too early may miss content that appears later.

Save multiple pages. If a directory spans multiple pages (pagination), save each page separately and upload them all in one batch. The browser extension can follow links to crawl up to 25 same-site pages automatically, which may be faster for paginated directories.

Webpage Loading vs File Upload

Email Extractor has two separate workflows:

  • Text and files (used in this guide): uploads files from your computer and processes them in your browser. No data leaves your machine.
  • Webpage loading: enters a URL and the server fetches the page. This is a different workflow that sends the URL to the server.

Webpage and file results do not automatically merge. If you use both methods, download each set of results separately and merge them if needed.

Troubleshooting

No addresses found

  • Open the saved HTML file in a text editor and search for @. If no @ symbols appear, the page either had no email addresses or they were loaded by JavaScript.
  • Try the browser extension on the live page instead. It scans the rendered page, including JavaScript-generated content.
  • Check that the file is actually HTML. A file with a .html extension that contains PDF data or another binary format will not produce results.

Unexpected addresses from page elements

  • Saved HTML files sometimes include addresses from navigation elements, site-wide footers or cookie consent scripts that appear on every page. These are legitimate addresses in the source code.
  • CMS-generated HTML may include administrator or system addresses in meta tags or comments. Review the results and remove any addresses that are not relevant to your purpose.

Source review

When uploading multiple HTML files in one batch, download results as CSV with sources to trace each address to its source file. This helps when extracting from several pages of the same website.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)