Building an Email Validation Pipeline: From Raw Data to Clean Lists
On this page
What Is a Validation Pipeline?
A validation pipeline is a series of automated steps that take raw email data and produce a clean, verified, ready-to-use list. Instead of manually uploading CSVs to different tools and downloading results, the pipeline runs end to end with minimal human intervention.
Why build one:
- Consistency. Every email that enters your system goes through the same checks.
- Speed. New data is cleaned within hours, not days.
- Scale. The pipeline handles 1,000 or 1,000,000 addresses the same way.
- Quality. Catch problems before they reach your CRM, email platform or sales team.
Pipeline Stages
Stage 1: Ingestion
Collect email data from all sources into a single staging area.
Sources to connect:
- Form submissions (website signups, lead magnets, event registrations).
- File uploads (CSV, XLSX, PDF imports from partners, data providers, purchased lists).
- CRM exports (periodic full exports or incremental syncs).
- API feeds (data provider APIs, marketing platform APIs).
- Manual entry (sales rep input, business card data).
Implementation:
- For file-based sources, extract email addresses using Email Extractor or a programmatic extraction step (regex or a parsing library for structured formats).
- For API sources, build connectors that pull new contacts on a schedule.
- For form submissions, capture directly from the form processor webhook.
- Write all ingested records to a staging table or file with: email, source, timestamp and any associated metadata.
Stage 2: Normalisation
Standardise email formatting before validation.
Steps:
- Trim whitespace. Remove leading, trailing and internal spaces.
- Lowercase. Convert the entire address to lowercase.
- Strip plus-addressing. Remove +tag portions (user+tag@example.com becomes user@example.com). Optional, depending on your use case.
- Fix common typos. Correct known domain misspellings:
- gnail.com to gmail.com
- yaho.com to yahoo.com
- outlok.com to outlook.com
- hotmal.com to hotmail.com
- Remove invalid characters. Strip characters that are not allowed in standard email addresses.
- Handle encoding. Convert HTML entities, URL encoding and unicode normalisation.
import re
def normalise_email(email):
if not email:
return None
# Trim and lowercase
email = email.strip().lower()
# Remove spaces within the address
email = email.replace(' ', '')
# Fix common domain typos
domain_fixes = {
'gnail.com': 'gmail.com',
'gmial.com': 'gmail.com',
'gmal.com': 'gmail.com',
'gamil.com': 'gmail.com',
'yaho.com': 'yahoo.com',
'yahooo.com': 'yahoo.com',
'hotmal.com': 'hotmail.com',
'hotmial.com': 'hotmail.com',
'outlok.com': 'outlook.com',
'outllook.com': 'outlook.com',
}
parts = email.split('@')
if len(parts) != 2:
return None
local, domain = parts
domain = domain_fixes.get(domain, domain)
return f"{local}@{domain}"
See Normalizing Email Formats.
Stage 3: Deduplication
Remove duplicate entries before spending money on verification.
Deduplication levels:
- Exact match. Same normalised address appears multiple times.
- Cross-source. Same address from multiple sources (keep the richest record).
- Gmail dot variants. For @gmail.com addresses, j.doe and jdoe deliver to the same mailbox.
Implementation:
- Group by normalised email address.
- When duplicates exist, keep the record with the most complete metadata, or the most recent.
- Track which sources contributed each address (useful for source quality analysis).
See How Email Deduplication Works.
Stage 4: Syntax validation
Check whether the email address follows valid syntax before making any network calls. This is free and instant.
Checks:
- Contains exactly one @ sign.
- Local part is not empty.
- Domain part is not empty.
- Domain contains at least one dot.
- TLD is at least 2 characters.
- No consecutive dots in the domain.
- No prohibited characters.
import re
def is_valid_syntax(email):
pattern = r'^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$'
return bool(re.match(pattern, email))
Action: Remove addresses that fail syntax validation. They will always bounce.
Stage 5: DNS/MX validation
Check whether the domain has mail servers configured to receive email.
How it works:
- Extract the domain from the email address.
- Query DNS for MX (Mail Exchange) records.
- If MX records exist, the domain can receive email.
- If no MX records, check for an A record (some domains accept email without MX records).
- If neither exists, the domain cannot receive email.
import dns.resolver
def has_mx_records(domain):
try:
records = dns.resolver.resolve(domain, 'MX')
return len(records) > 0
except (dns.resolver.NoAnswer,
dns.resolver.NXDOMAIN,
dns.resolver.NoNameservers,
dns.exception.Timeout):
return False
Action: Remove addresses with no MX or A records. The domain does not accept email.
Performance tip: Cache MX lookup results by domain. Multiple addresses at the same domain need only one DNS query.
See MX Records Explained.
Stage 6: Disposable and role-based detection
Identify addresses that are valid but low-quality for outreach.
Disposable addresses: Temporary email services (Guerrilla Mail, Temp Mail, Mailinator). Addresses expire within hours or days.
Role-based addresses: Shared inboxes (info@, admin@, sales@, support@, webmaster@). Not tied to an individual.
Implementation: Maintain or subscribe to lists of known disposable email domains and role-based local parts.
DISPOSABLE_DOMAINS = {'guerrillamail.com', 'tempmail.com',
'mailinator.com', 'throwaway.email', ...}
ROLE_PREFIXES = {'info', 'admin', 'sales', 'support', 'help',
'contact', 'webmaster', 'postmaster', 'abuse',
'noreply', 'no-reply', 'billing', 'marketing'}
def classify_email(email):
local, domain = email.split('@')
flags = []
if domain in DISPOSABLE_DOMAINS:
flags.append('disposable')
if local in ROLE_PREFIXES:
flags.append('role_based')
return flags
Action: Flag (do not necessarily remove) disposable and role-based addresses. The decision to include or exclude depends on the use case.
See Disposable Email Addresses and Role-Based Addresses.
Stage 7: SMTP verification
Connect to the recipient's mail server and check whether the specific mailbox exists. This is the most definitive check but also the slowest and most resource-intensive.
How it works:
- Connect to the MX server on port 25.
- Send HELO/EHLO command.
- Send MAIL FROM command.
- Send RCPT TO command with the email address.
- The server responds with acceptance (250) or rejection (550).
- Disconnect without sending a message.
Limitations:
- Catch-all domains. Some servers accept all addresses regardless. SMTP verification returns "valid" even for non-existent mailboxes.
- Greylisting. Some servers temporarily reject the first connection attempt.
- Rate limiting. Servers may block or delay repeated verification attempts.
- Privacy-conscious servers. Some servers do not reveal whether a mailbox exists (they accept all RCPT TO commands).
Action: For most teams, using a third-party verification API (which handles SMTP verification at scale, manages IP reputation and handles retries) is more practical than building SMTP verification in-house.
See SMTP Email Verification Explained.
Stage 8: Third-party verification API
For production pipelines, send addresses through a verification API that handles all of the above checks plus additional signals.
Integration:
def verify_batch(emails, api_key):
"""Send batch to verification service."""
response = requests.post(
'https://api.verificationservice.example.com/v1/batch',
json={'emails': emails},
headers={'Authorization': f'Bearer {api_key}'},
timeout=30
)
return response.json()['batch_id']
# Submit, wait, retrieve
batch_id = verify_batch(email_list, API_KEY)
results = poll_for_results(batch_id)
See Email Validation API Integration and Email Verification Service Comparison.
Stage 9: Scoring and categorisation
Assign a quality score to each address based on the validation results.
def score_email(result):
score = 0
if result['valid']:
score += 50
if not result['disposable']:
score += 15
if not result['role_based']:
score += 10
if not result['free_provider']:
score += 10
if not result['catch_all']:
score += 15
return score
Categorisation:
| Score | Category | Action |
|---|---|---|
| 80-100 | High quality | Add to active outreach lists |
| 50-79 | Medium quality | Add with caution, monitor bounce rates |
| 20-49 | Low quality | Use for low-risk campaigns only |
| 0-19 | Poor quality | Suppress or remove |
Stage 10: Output and routing
Route validated addresses to the appropriate destination:
- High-quality addresses go to the CRM and outreach sequences.
- Medium-quality addresses go to nurture lists.
- Invalid addresses go to a rejection log for analysis.
- Disposable and role-based addresses go to a flagged list for manual review.
Automation and Scheduling
Real-time pipeline
For form submissions and API integrations, run validation in real time:
- Form submits email address.
- Pipeline runs stages 2-8 in under 3 seconds.
- Valid address is accepted; invalid is rejected with a user-friendly error.
Batch pipeline
For file uploads and periodic cleaning, run on a schedule:
- New files land in an ingestion folder.
- A scheduled job runs the pipeline every hour (or daily).
- Results are delivered to the output destination.
Periodic re-verification
Schedule re-verification of your entire database every 3-6 months:
- Export all active contacts.
- Run through stages 5-9.
- Flag newly invalid addresses for removal.
- Report on list decay rate.
See Email List Decay Explained.
Monitoring
Track pipeline health with these metrics:
- Volume. How many addresses processed per day/week.
- Pass rate. What percentage of addresses pass validation.
- Source quality. Pass rate broken down by source (which sources produce the cleanest data).
- Processing time. How long the pipeline takes (important for real-time use cases).
- API costs. Verification API spend per period.
- Bounce rate (post-pipeline). Addresses that passed validation but still bounced when emailed. This is your pipeline's false positive rate.