Regex for Email Extraction: Patterns, Examples and Limitations
On this page
What Regex Does for Email Extraction
Regular expressions (regex) are patterns that match text. When applied to a body of text, a regex pattern can find every string that looks like an email address and pull it out.
This is the foundation of how email extraction works: scan text, match patterns that follow the format local-part@domain, and return the matches.
The Basic Pattern
The simplest email regex:
[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}
What each part matches:
| Part | Matches | Example |
|---|---|---|
[a-zA-Z0-9._%+-]+ |
Local part (before @) | john.doe, info, user+tag |
@ |
Literal @ sign | @ |
[a-zA-Z0-9.-]+ |
Domain name | example, mail.company |
\. |
Literal dot | . |
[a-zA-Z]{2,} |
Top-level domain | com, org, co.uk |
What it catches: Most standard business and personal email addresses.
What it misses or gets wrong: Addresses with unusual but valid characters, internationalised domains, very long TLDs, and edge cases covered below.
Language Examples
Python
import re
text = """
Contact us at info@example.com or sales@example.org.
You can also reach john.doe+newsletter@example.net.
"""
pattern = r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}'
emails = re.findall(pattern, text)
print(emails)
# ['info@example.com', 'sales@example.org',
# 'john.doe+newsletter@example.net']
JavaScript
const text = `
Contact us at info@example.com or sales@example.org.
You can also reach john.doe+newsletter@example.net.
`;
const pattern = /[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}/g;
const emails = text.match(pattern);
console.log(emails);
// ['info@example.com', 'sales@example.org',
// 'john.doe+newsletter@example.net']
Bash (grep)
grep -oE '[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}' input.txt
Google Sheets
=REGEXEXTRACT(A1,"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}")
Note: REGEXEXTRACT returns only the first match per cell. For multiple emails in one cell, you would need a custom function or Google Apps Script.
Excel (Microsoft 365)
Excel does not have a built-in regex function. Options include:
- VBA with the
VBScript.RegExpobject. - Power Query's Text.Select or custom M functions.
- The FILTERXML trick (limited and fragile).
For bulk extraction from spreadsheets, uploading the file to Email Extractor is simpler and handles XLSX, XLS, ODS and CSV formats.
Advanced Patterns
Handling subdomains
The basic pattern already handles subdomains (user@mail.example.com) because [a-zA-Z0-9.-]+ matches dots within the domain.
Case-insensitive matching
Email addresses are case-insensitive in the local part per RFC 5321 (though some servers treat them as case-sensitive in practice). Add the case-insensitive flag:
Python: re.findall(pattern, text, re.IGNORECASE)
JavaScript: /pattern/gi
Bash: grep -oEi 'pattern' input.txt
Avoiding partial matches
Without word boundaries, the regex might match inside URLs or other strings:
Visit https://user@example.com/page
Add word boundaries or negative lookbehinds to prevent this:
(?<![/=:])([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})
Extracting from HTML
HTML email addresses are often wrapped in mailto links:
<a href="mailto:info@example.com">Contact us</a>
A regex for mailto links:
mailto:([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})
Handling obfuscated addresses
Some websites obfuscate email addresses to prevent scraping:
| Obfuscation | Example |
|---|---|
| [at] replacement | user [at] example [dot] com |
| HTML entities | user@example.com |
| Reversed text | moc.elpmaxe@resu |
| JavaScript assembly | Building the address dynamically in JS |
Regex can handle simple replacements:
# Handle [at] and [dot] variations
text = text.replace(' [at] ', '@').replace('[at]', '@')
text = text.replace(' [dot] ', '.').replace('[dot]', '.')
text = text.replace(' (at) ', '@').replace('(at)', '@')
text = text.replace(' (dot) ', '.').replace('(dot)', '.')
# Then apply standard regex
For JavaScript-assembled addresses and complex obfuscation, regex alone is not sufficient. You need a browser that executes JavaScript. See How Websites Block Email Extraction.
Edge Cases
Valid addresses that basic regex misses
The RFC 5321 specification allows characters in email addresses that most regex patterns do not match:
| Valid but unusual | Example |
|---|---|
| Quoted local part | "john doe"@example.com |
| Special characters | user!tag@example.com |
| Consecutive dots | user..name@example.com (valid per RFC, rejected by most servers) |
| IP address domain | user@[192.168.1.1] |
| Long TLDs | user@example.photography |
| Internationalised domains | user@example.xn--e1afmapc (Punycode) |
| Unicode local part | user@example.com (some newer servers) |
Practical guidance: Most of these are rare in real-world data. The basic pattern catches 99%+ of addresses you will encounter. Building a regex that matches every RFC-valid address creates more false positives than it prevents false negatives.
False positives
The basic pattern can match strings that look like email addresses but are not:
| False positive | Why it matches |
|---|---|
| filename@2x.png | Looks like user@domain.tld |
| version@1.0.0 | @ followed by dotted numbers |
| user@localhost | No dot-separated TLD (filtered by the pattern) |
| SHA@commit.abc | Random text with @ |
Mitigation:
- Require a minimum TLD length of 2 characters (already in the basic pattern).
- Maintain a list of known non-email TLDs to filter against.
- Post-process matches with a domain existence check (MX record lookup).
Addresses split across lines
In some documents, email addresses wrap across line breaks:
Contact john.doe@
example.com for details.
Standard regex does not match across line breaks by default. Options:
- Pre-process the text to join lines.
- Use the
re.DOTALLorre.MULTILINEflag (depending on the pattern). - Use a tool like Email Extractor that handles this automatically.
When Regex Is Not Enough
Binary file formats
Regex works on text. Binary file formats like DOCX (zipped XML), XLSX (zipped XML), PDF (binary with text streams) and MSG (Microsoft Outlook format) require parsing the file format first, then applying regex to the extracted text.
DOCX: Unzip, read the XML content from word/document.xml, strip XML tags, apply regex.
XLSX: Unzip, read shared strings from xl/sharedStrings.xml and cell values, apply regex.
PDF: Use a PDF text extraction library (pdftotext, PyPDF2, pdfplumber), then apply regex to the extracted text. Note: scanned PDFs contain images, not text. Text extraction will return nothing unless OCR is applied.
EML/MSG: Parse the email format to extract headers and body text, then apply regex.
Email Extractor handles all 19 of these file formats in the browser. You upload the file, and the tool parses the format, extracts text and applies pattern matching without needing to write code.
JavaScript-rendered content
Websites that load email addresses via JavaScript will not yield those addresses to a simple text scrape. The HTML source contains the JavaScript code, not the rendered email address. You need a headless browser (Puppeteer, Playwright) to render the page and then extract from the rendered DOM.
Image-based text
Email addresses in images (screenshots, scanned documents, business card photos) require OCR (Optical Character Recognition) before regex can be applied. Tools like Tesseract, Google Vision API or AWS Textract convert image text to machine-readable text.
Encoded addresses
Addresses encoded as HTML entities, URL-encoded strings or Base64 require decoding before regex matching:
import html
text = html.unescape("info@example.com")
# Result: "info@example.com"
Performance Considerations
Large files
For files over 10 MB, regex performance matters:
- Compile the pattern once. In Python, use
re.compile(pattern)and reuse the compiled object. - Read in chunks. For very large files, process the file in chunks rather than loading the entire content into memory.
- Avoid catastrophic backtracking. Poorly written regex patterns can take exponential time on certain inputs. The basic email pattern above does not have this problem, but more complex patterns can.
Deduplication after extraction
Regex extraction often returns duplicates (the same address appearing multiple times in a document). Deduplicate by converting to a set:
emails = list(set(re.findall(pattern, text)))
For case-insensitive deduplication:
seen = set()
unique = []
for email in re.findall(pattern, text):
lower = email.lower()
if lower not in seen:
seen.add(lower)
unique.append(email)
Regex vs Dedicated Tools
| Factor | Regex (DIY) | Email Extractor |
|---|---|---|
| Text files | Works well | Works well |
| Binary formats (DOCX, XLSX, PDF) | Requires parsing library | Handled automatically |
| Deduplication | Manual post-processing | Automatic |
| Edge cases | Manual handling | Handled automatically |
| Setup time | Minutes to hours | Seconds |
| Programming required | Yes | No |
For one-off extraction from a plain text file, regex is quick and effective. For extracting from multiple file formats, deduplicating across sources, and handling edge cases, a dedicated tool like Email Extractor saves significant time.