Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Regex for Email Extraction: Patterns, Examples and Limitations

On this page

What Regex Does for Email Extraction

Regular expressions (regex) are patterns that match text. When applied to a body of text, a regex pattern can find every string that looks like an email address and pull it out.

This is the foundation of how email extraction works: scan text, match patterns that follow the format local-part@domain, and return the matches.

The Basic Pattern

The simplest email regex:

[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}

What each part matches:

Part Matches Example
[a-zA-Z0-9._%+-]+ Local part (before @) john.doe, info, user+tag
@ Literal @ sign @
[a-zA-Z0-9.-]+ Domain name example, mail.company
\. Literal dot .
[a-zA-Z]{2,} Top-level domain com, org, co.uk

What it catches: Most standard business and personal email addresses.

What it misses or gets wrong: Addresses with unusual but valid characters, internationalised domains, very long TLDs, and edge cases covered below.

Language Examples

Python

import re

text = """
Contact us at info@example.com or sales@example.org.
You can also reach john.doe+newsletter@example.net.
"""

pattern = r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}'
emails = re.findall(pattern, text)

print(emails)
# ['info@example.com', 'sales@example.org',
#  'john.doe+newsletter@example.net']

JavaScript

const text = `
Contact us at info@example.com or sales@example.org.
You can also reach john.doe+newsletter@example.net.
`;

const pattern = /[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}/g;
const emails = text.match(pattern);

console.log(emails);
// ['info@example.com', 'sales@example.org',
//  'john.doe+newsletter@example.net']

Bash (grep)

grep -oE '[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}' input.txt

Google Sheets

=REGEXEXTRACT(A1,"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}")

Note: REGEXEXTRACT returns only the first match per cell. For multiple emails in one cell, you would need a custom function or Google Apps Script.

Excel (Microsoft 365)

Excel does not have a built-in regex function. Options include:

  • VBA with the VBScript.RegExp object.
  • Power Query's Text.Select or custom M functions.
  • The FILTERXML trick (limited and fragile).

For bulk extraction from spreadsheets, uploading the file to Email Extractor is simpler and handles XLSX, XLS, ODS and CSV formats.

Advanced Patterns

Handling subdomains

The basic pattern already handles subdomains (user@mail.example.com) because [a-zA-Z0-9.-]+ matches dots within the domain.

Case-insensitive matching

Email addresses are case-insensitive in the local part per RFC 5321 (though some servers treat them as case-sensitive in practice). Add the case-insensitive flag:

Python: re.findall(pattern, text, re.IGNORECASE)

JavaScript: /pattern/gi

Bash: grep -oEi 'pattern' input.txt

Avoiding partial matches

Without word boundaries, the regex might match inside URLs or other strings:

Visit https://user@example.com/page

Add word boundaries or negative lookbehinds to prevent this:

(?<![/=:])([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})

Extracting from HTML

HTML email addresses are often wrapped in mailto links:

<a href="mailto:info@example.com">Contact us</a>

A regex for mailto links:

mailto:([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})

Handling obfuscated addresses

Some websites obfuscate email addresses to prevent scraping:

Obfuscation Example
[at] replacement user [at] example [dot] com
HTML entities user@example.com
Reversed text moc.elpmaxe@resu
JavaScript assembly Building the address dynamically in JS

Regex can handle simple replacements:

# Handle [at] and [dot] variations
text = text.replace(' [at] ', '@').replace('[at]', '@')
text = text.replace(' [dot] ', '.').replace('[dot]', '.')
text = text.replace(' (at) ', '@').replace('(at)', '@')
text = text.replace(' (dot) ', '.').replace('(dot)', '.')
# Then apply standard regex

For JavaScript-assembled addresses and complex obfuscation, regex alone is not sufficient. You need a browser that executes JavaScript. See How Websites Block Email Extraction.

Edge Cases

Valid addresses that basic regex misses

The RFC 5321 specification allows characters in email addresses that most regex patterns do not match:

Valid but unusual Example
Quoted local part "john doe"@example.com
Special characters user!tag@example.com
Consecutive dots user..name@example.com (valid per RFC, rejected by most servers)
IP address domain user@[192.168.1.1]
Long TLDs user@example.photography
Internationalised domains user@example.xn--e1afmapc (Punycode)
Unicode local part user@example.com (some newer servers)

Practical guidance: Most of these are rare in real-world data. The basic pattern catches 99%+ of addresses you will encounter. Building a regex that matches every RFC-valid address creates more false positives than it prevents false negatives.

False positives

The basic pattern can match strings that look like email addresses but are not:

False positive Why it matches
filename@2x.png Looks like user@domain.tld
version@1.0.0 @ followed by dotted numbers
user@localhost No dot-separated TLD (filtered by the pattern)
SHA@commit.abc Random text with @

Mitigation:

  • Require a minimum TLD length of 2 characters (already in the basic pattern).
  • Maintain a list of known non-email TLDs to filter against.
  • Post-process matches with a domain existence check (MX record lookup).

Addresses split across lines

In some documents, email addresses wrap across line breaks:

Contact john.doe@
example.com for details.

Standard regex does not match across line breaks by default. Options:

  • Pre-process the text to join lines.
  • Use the re.DOTALL or re.MULTILINE flag (depending on the pattern).
  • Use a tool like Email Extractor that handles this automatically.

When Regex Is Not Enough

Binary file formats

Regex works on text. Binary file formats like DOCX (zipped XML), XLSX (zipped XML), PDF (binary with text streams) and MSG (Microsoft Outlook format) require parsing the file format first, then applying regex to the extracted text.

DOCX: Unzip, read the XML content from word/document.xml, strip XML tags, apply regex.

XLSX: Unzip, read shared strings from xl/sharedStrings.xml and cell values, apply regex.

PDF: Use a PDF text extraction library (pdftotext, PyPDF2, pdfplumber), then apply regex to the extracted text. Note: scanned PDFs contain images, not text. Text extraction will return nothing unless OCR is applied.

EML/MSG: Parse the email format to extract headers and body text, then apply regex.

Email Extractor handles all 19 of these file formats in the browser. You upload the file, and the tool parses the format, extracts text and applies pattern matching without needing to write code.

JavaScript-rendered content

Websites that load email addresses via JavaScript will not yield those addresses to a simple text scrape. The HTML source contains the JavaScript code, not the rendered email address. You need a headless browser (Puppeteer, Playwright) to render the page and then extract from the rendered DOM.

Image-based text

Email addresses in images (screenshots, scanned documents, business card photos) require OCR (Optical Character Recognition) before regex can be applied. Tools like Tesseract, Google Vision API or AWS Textract convert image text to machine-readable text.

Encoded addresses

Addresses encoded as HTML entities, URL-encoded strings or Base64 require decoding before regex matching:

import html
text = html.unescape("info&#64;example&#46;com")
# Result: "info@example.com"

Performance Considerations

Large files

For files over 10 MB, regex performance matters:

  • Compile the pattern once. In Python, use re.compile(pattern) and reuse the compiled object.
  • Read in chunks. For very large files, process the file in chunks rather than loading the entire content into memory.
  • Avoid catastrophic backtracking. Poorly written regex patterns can take exponential time on certain inputs. The basic email pattern above does not have this problem, but more complex patterns can.

Deduplication after extraction

Regex extraction often returns duplicates (the same address appearing multiple times in a document). Deduplicate by converting to a set:

emails = list(set(re.findall(pattern, text)))

For case-insensitive deduplication:

seen = set()
unique = []
for email in re.findall(pattern, text):
    lower = email.lower()
    if lower not in seen:
        seen.add(lower)
        unique.append(email)

Regex vs Dedicated Tools

Factor Regex (DIY) Email Extractor
Text files Works well Works well
Binary formats (DOCX, XLSX, PDF) Requires parsing library Handled automatically
Deduplication Manual post-processing Automatic
Edge cases Manual handling Handled automatically
Setup time Minutes to hours Seconds
Programming required Yes No

For one-off extraction from a plain text file, regex is quick and effective. For extracting from multiple file formats, deduplicating across sources, and handling edge cases, a dedicated tool like Email Extractor saves significant time.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)