Email Obfuscation: Why Some Addresses Can't Be Extracted
On this page
What Email Obfuscation Is
Email obfuscation is the practice of hiding email addresses from automated tools while keeping them readable (or recoverable) by humans. Website owners and document creators use obfuscation to reduce spam and unwanted messages.
When you extract email addresses from a file or web page, obfuscated addresses may not be found because they do not match the standard email pattern (local-part@domain.tld) that extraction tools look for.
Understanding obfuscation helps you recognise why an expected address might be missing from your extraction results and what alternatives are available.
Common Techniques
Text substitution
The most basic form replaces the @ symbol or dot with words:
contact [at] example [dot] comcontact(at)example.comcontact AT example DOT com
These are readable by humans but do not match the standard email address pattern. Email Extractor will not find them in this form.
HTML entity encoding
On web pages, the @ symbol can be written as its HTML entity (@ or @). The browser renders it normally, but the raw HTML source contains the encoded version.
When you save a web page as HTML and upload it, Email Extractor reads the source code. Whether encoded entities are found depends on how the extraction processes the HTML. Some encoding methods are transparent to extraction; others are not.
JavaScript assembly
Some websites use JavaScript to assemble email addresses from parts at render time:
var user = "contact";
var domain = "example.com";
document.write(user + "@" + domain);
The address never appears complete in the HTML source. When you save the page as HTML, the JavaScript code is saved but not executed. Since Email Extractor does not execute JavaScript in uploaded files, the address is not found.
The browser extension scans the rendered page after JavaScript has executed, so it may find addresses that are assembled by JavaScript.
CSS direction tricks
Some websites display email addresses using CSS to reverse the text direction:
<span style="unicode-bidi: bidi-override; direction: rtl;">moc.elpmaxe@tcatnoc</span>
The browser renders this as contact@example.com, but the underlying text is reversed. Extraction from the saved HTML source finds the reversed version, which is not a valid address.
Image-based addresses
Some websites render email addresses as images instead of text. The address is visible on the page but exists only as pixels. No text-based extraction tool can read it.
Cloudflare email protection
Websites using Cloudflare's email address obfuscation replace visible addresses with encoded strings. The original address is decoded by a JavaScript snippet at render time. Saved HTML files contain the encoded version, not the original address.
Mailto link obfuscation
Some websites use JavaScript to construct mailto: links dynamically, or encode the entire link so it only works when clicked in a browser. Extraction from saved HTML will not find these.
What Email Extractor Can and Cannot Find
Will find:
- Plain text email addresses in any supported file type.
- Addresses in
mailto:links in saved HTML files (when the link is written in the source). - Addresses in HTML comments and meta tags.
- Addresses encoded with simple HTML entities that resolve to standard characters during parsing.
Will not find:
- Addresses assembled by JavaScript (in saved HTML files).
- Addresses rendered as images.
- Addresses using text substitution (
[at],[dot]). - Addresses reversed by CSS.
- Addresses protected by Cloudflare email obfuscation or similar services (in saved HTML files).
May find (depending on method):
- The browser extension scans the rendered page, which means it can find JavaScript-assembled addresses and Cloudflare-decoded addresses on live pages. It does not read image-based addresses.
What to Do When Addresses Are Hidden
Try the browser extension
If a web page shows email addresses visually but saving and uploading the HTML produces no results, try the browser extension on the live page. It scans what the browser has rendered, including JavaScript output.
Check the page source
Right-click the page in your browser and select "View Page Source" or press Ctrl+U. Search for @ in the source. If you see the address in plain text in the source, the saved HTML file should contain it. If you see JavaScript code or encoded strings instead, the address is obfuscated.
Look for alternative pages
The same website may list contact information on multiple pages. A team page might obfuscate addresses while a press page lists them in plain text. Try extracting from different pages.
Use the contact form
If a website has deliberately obfuscated its email addresses, the preferred contact method may be a form rather than direct email. The obfuscation is a signal that the site owner prefers to control how they receive messages.
Check other sources
The same email address may appear in unobfuscated form elsewhere: in a PDF document, a conference attendee list, a business directory, a professional network profile, or a public filing. Search for the person or organisation name in your existing files before giving up.
Why Obfuscation Exists
Website owners obfuscate email addresses primarily to reduce automated harvesting by spambots. Publishing an email address in plain text on a public web page exposes it to automated scraping, which leads to spam.
Obfuscation is a defensive measure by the site owner. Its presence indicates that the site owner has chosen to limit automated access to their contact information.