Regex Patterns for Email Extraction: A Practical Reference
On this page
A pattern finds candidates, not working mailboxes
A regular expression can collect email-like strings from plain text. It does not establish whether a domain accepts mail, whether a mailbox exists or whether you have permission to contact it.
The examples below are deliberately small ASCII candidate detectors. They are useful for learning and controlled text inputs, not complete email syntax validators. For document readers and source tracking without writing code, use Email Extractor.
Understand the basic pattern
[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}
The first group finds a limited set of local-part characters. The @ separates it from a domain, and the final group requires at least two ASCII letters after a dot. This can match subdomains and the whole address sam@department.example.co.uk.
The pattern also accepts malformed candidates such as consecutive local-part dots. It can match only part of an address containing unsupported characters. For example, #team@example.org becomes team@example.org, which is a different address. Do not use this limited pattern when preserving such inputs matters.
Python example
import re
from pathlib import Path
pattern = r"[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}"
text = Path("file.txt").read_text(encoding="utf-8")
for address in sorted(set(re.findall(pattern, text))):
print(address)
This reads UTF-8 text, finds candidates, removes exact repeated strings and sorts the output. Its duplicate comparison is case-sensitive; that is a deliberate property of this sample, not a universal rule for mailbox identity. Python's regular expression documentation explains findall and pattern behavior.
JavaScript example
const text = "Contact sam@example.com or help@example.org";
const pattern = /[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}/g;
const addresses = [...new Set(text.match(pattern) || [])].sort();
console.log(addresses);
The g flag collects multiple matches. The empty-array fallback handles text with no candidates. These examples do not decode HTML entities, JSON escapes or mailto: percent encoding before matching.
Check boundary and syntax cases
| Input | Basic pattern result | What to review |
|---|---|---|
sam@example.com |
sam@example.com |
A candidate, not a verified mailbox |
sam+news@example.org |
Same address | Preserve the plus tag |
#team@example.org |
team@example.org |
Truncated local part; unsuitable output |
a..b@example.com |
Same string | Consecutive dots need syntax review |
sam@department.example.co.uk |
Same address | Subdomains are included |
sam@example.com repeated |
One result after deduplication | Exact-string duplicate only |
A trailing sentence period can be punctuation, while a leading # can be part of a real local part. Context and boundaries matter; removing every unfamiliar character is not a safe cleanup rule.
More characters do not make a complete validator
The email message specification permits additional unquoted local-part characters, including # and *, and describes quoted local parts. Replacing a character class with \w does not implement all those rules or validate an internationalized domain.
These examples exclude quoted local parts, address literals and internationalized addresses. Use a maintained parser with an explicitly documented scope when those formats are required. Binary PDF, XLSX and MSG files also need their format decoded before text matching.
Before using any extraction script on production data, test positive, negative and truncation cases drawn from your own sources. Keep the original strings for review. Read email validation after extraction if you also need domain or mailbox checks.