Email Verification at Scale: A Guide for High-Volume Senders
By Email ExtractorPublished 8 min read
On this page
When Standard Verification Is Not Enough
Most email verification guidance assumes lists of a few thousand contacts. At 100,000+ contacts, new challenges appear:
Challenge
At small scale
At high volume
Cost
A few dollars per batch
Thousands of dollars per verification cycle
Processing time
Minutes
Hours to days
API rate limits
Rarely an issue
Frequently hit; requires queuing
Accuracy trade-offs
Verify everything
Must prioritise where verification adds the most value
Data freshness
Verify before each send
Must balance freshness against cost
Infrastructure
Single API call
Pipeline with queuing, retries and error handling
Vendor dependence
Low risk
Single vendor outage can block sending
Verification Architecture for Scale
Pipeline design
A high-volume verification pipeline has several stages:
Stage
Purpose
What happens
1. Syntax check
Remove obviously invalid addresses
Regex or library-based check; no API cost
2. Domain check
Remove addresses at non-existent domains
DNS MX lookup; no API cost
3. Disposable domain check
Remove temporary email addresses
Check against disposable domain lists; no API cost
4. Role-based check
Flag or remove role addresses (info@, support@)
Pattern matching; no API cost
5. Deduplication
Remove duplicates before paying for verification
Exact match and normalisation; no API cost
6. SMTP verification
Check if the mailbox exists
API call to verification service; primary cost
7. Catch-all detection
Identify domains that accept all addresses
Part of SMTP verification; affects confidence level
8. Result classification
Categorise results for action
Processing logic; no API cost
Running stages 1-5 before stage 6 reduces the number of addresses sent to the paid verification service. On a typical list, this pre-filtering removes 5-15% of addresses before any API cost is incurred.
Pre-filtering with Email Extractor
Before sending a list to a verification service, extract and deduplicate the addresses:
Upload your source files (CSV, XLSX, TXT or other supported formats) to Email Extractor.
The tool performs case-insensitive deduplication automatically.
Download the deduplicated list as CSV.
Run syntax and domain checks locally.
Send only the survivors to your verification API.
This saves verification costs by removing duplicates and allowing you to pre-filter before the paid step.
Code: Local pre-filtering
import re
import dns.resolver
def syntax_check(email):
"""Basic syntax validation. Returns True if the format looks valid."""
pattern = r'^[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}$'
return bool(re.match(pattern, email))
def domain_has_mx(domain, cache={}):
"""Check if the domain has MX records. Cache results."""
if domain in cache:
return cache[domain]
try:
dns.resolver.resolve(domain, 'MX')
cache[domain] = True
except (dns.resolver.NXDOMAIN, dns.resolver.NoAnswer,
dns.resolver.NoNameservers, dns.resolver.Timeout):
cache[domain] = False
return cache[domain]
# Known disposable email domains (partial list; use a maintained list in production)
DISPOSABLE_DOMAINS = {
'tempmail.com', 'throwaway.email', 'guerrillamail.com',
'mailinator.com', 'yopmail.com'
# In production, use a maintained list with thousands of domains
}
ROLE_PREFIXES = {
'info', 'support', 'admin', 'sales', 'contact', 'help',
'billing', 'abuse', 'postmaster', 'webmaster', 'noreply',
'no-reply', 'marketing', 'press', 'media', 'office'
}
def pre_filter(emails):
"""Run free pre-filtering before paid verification."""
results = {
'valid': [],
'invalid_syntax': [],
'no_mx': [],
'disposable': [],
'role_based': []
}
seen = set()
for email in emails:
email = email.strip().lower()
# Deduplicate
if email in seen:
continue
seen.add(email)
# Syntax check
if not syntax_check(email):
results['invalid_syntax'].append(email)
continue
local_part, domain = email.rsplit('@', 1)
# Disposable domain check
if domain in DISPOSABLE_DOMAINS:
results['disposable'].append(email)
continue
# MX record check
if not domain_has_mx(domain):
results['no_mx'].append(email)
continue
# Role-based check (flag but do not remove)
if local_part in ROLE_PREFIXES:
results['role_based'].append(email)
results['valid'].append(email)
return results
Cost Management
Verification pricing at scale
Volume tier
Typical price per verification
Monthly cost at 500K verifications
Pay-as-you-go
$0.005-0.01
$2,500-5,000
Volume plan (100K+)
$0.003-0.006
$1,500-3,000
Enterprise plan (1M+)
$0.001-0.003
$500-1,500
Self-hosted (infrastructure cost)
Variable
Depends on infrastructure
Strategies to reduce verification costs
Strategy
Savings
Trade-off
Pre-filter before verification
5-15% fewer API calls
Requires local processing
Verify only new additions
Only pay for new addresses
Existing addresses may have gone stale
Tiered verification frequency
Verify active contacts less often
Some stale addresses may slip through
Cache verification results
Avoid re-verifying the same address
Results go stale over time
Use multiple vendors strategically
Lower cost per verification
More complex integration
Verify on ingest, not before send
Spreads cost over time
Addresses may change between ingest and send
Verification frequency recommendations
Contact type
Verification frequency
Rationale
New additions
At time of ingest
Catch bad data before it enters your system
Active contacts (opened/clicked in 30 days)
Every 6 months
Low risk; recently engaged
Inactive contacts (no engagement in 30-90 days)
Every 3 months
Higher risk of going stale
Dormant contacts (no engagement in 90+ days)
Before each send
Highest risk; verify or remove
Re-engagement campaigns
Before sending
By definition, these addresses have not engaged recently
Purchased or rented lists
Before first use
Quality is unknown; expect high invalid rates
Handling Verification Results at Scale
Result categories and actions
Result
Meaning
Action
Valid
Mailbox exists and accepts mail
Safe to send
Invalid
Mailbox does not exist
Remove immediately
Catch-all
Domain accepts all addresses; individual validity unknown
Send with caution; monitor bounces
Unknown
Verification could not determine status (timeout, greylisting)
Retry later; send with caution if retry also fails
Disposable
Temporary/throwaway address
Remove or flag based on use case
Role-based
Generic address (info@, support@)
Flag; may be appropriate for some campaigns
Spam trap (suspected)
Some services flag potential spam traps
Remove immediately
Catch-all domain handling at scale
Catch-all domains are a particular challenge at high volume. When a domain accepts all addresses, verification cannot determine whether a specific mailbox exists:
Approach
When to use
Send to all catch-all results
When your sender reputation is strong and you can absorb some bounces
Send to catch-all only if other signals are positive
When you have engagement data or other validation
Skip catch-all addresses
When your sender reputation is fragile or you are warming up a new domain
Verify catch-all addresses with a secondary service
When the volume justifies the additional cost
Monitor catch-all bounce rates separately
Always; this tells you whether your catch-all strategy is working
Batch processing workflow
import time
import csv
from collections import defaultdict
def process_verification_results(results_file, output_dir):
"""Process bulk verification results and split into action files."""
actions = defaultdict(list)
with open(results_file, 'r') as f:
reader = csv.DictReader(f)
for row in reader:
email = row['email']
result = row['result'].lower()
reason = row.get('reason', '')
if result == 'valid':
actions['send'].append(row)
elif result == 'invalid':
actions['remove'].append(row)
elif result == 'catch_all':
actions['catch_all_review'].append(row)
elif result == 'unknown':
actions['retry'].append(row)
elif result == 'disposable':
actions['remove'].append(row)
elif result in ('role', 'role_based'):
actions['role_review'].append(row)
else:
actions['manual_review'].append(row)
# Write each action group to a separate file
for action, rows in actions.items():
output_file = f"{output_dir}/{action}.csv"
if rows:
with open(output_file, 'w', newline='') as f:
writer = csv.DictWriter(f, fieldnames=rows[0].keys())
writer.writeheader()
writer.writerows(rows)
print(f"{action}: {len(rows)} addresses")
return dict(actions)
Multi-Vendor Strategy
Why use multiple verification services
Reason
Explanation
Redundancy
If one service goes down, you can fall back to another
Accuracy
Different services have different strengths; cross-referencing improves accuracy
Cost optimisation
Use the cheapest service for bulk, a more accurate one for edge cases
Rate limit management
Spread load across multiple APIs
Catch-all handling
Some services are better at catch-all detection than others
Multi-vendor architecture
Tier
Service
Use case
Primary
High-volume, cost-effective service
Bulk verification of new lists
Secondary
High-accuracy service
Re-verify unknowns and catch-all addresses from primary
Tertiary
Specialist service
Spam trap detection, advanced catch-all analysis
Consensus-based verification
For critical sends, verify an address with multiple services and use consensus:
Service A result
Service B result
Action
Valid
Valid
Send with confidence
Valid
Invalid
Re-verify with service C; lean toward removing
Invalid
Invalid
Remove with confidence
Valid
Unknown
Send with monitoring
Unknown
Unknown
Retry later or remove if retries also fail
Catch-all
Catch-all
Both agree it is catch-all; apply catch-all strategy
Valid
Catch-all
Domain is likely catch-all; treat as catch-all
Operational Considerations
Monitoring verification health
Metric
What it tells you
Alert threshold
Invalid rate on new lists
Quality of your data sources
Above 20% suggests a bad source
Invalid rate on existing lists
Rate of list decay
Above 5% per quarter suggests verification is not frequent enough
Unknown rate
Verification service reliability
Above 10% suggests service issues
Catch-all rate
Proportion of unverifiable addresses
Track trend; increasing catch-all may require strategy change
Verification API response time
Service performance
Sustained increase may indicate throttling
Verification API error rate
Service reliability
Above 1% sustained warrants investigation
Post-send bounce rate
Accuracy of verification
Should be under 2% for verified lists
Data retention and compliance
Requirement
Implementation
Store verification results with timestamps
Know when each address was last verified
Maintain verification audit trail
Track which service verified each address and the result
Honour data deletion requests
GDPR and similar laws require the ability to delete contact records
Document verification processes
Demonstrate due diligence for compliance purposes
Separate verification data from PII where possible
Minimise data exposure in case of breach
Scaling Considerations
Scale
Architecture
Key challenges
Under 10K contacts
Manual batch upload to verification service
Minimal; straightforward
10K-100K contacts
API integration with queuing
Rate limits, cost management
100K-1M contacts
Pipeline with pre-filtering, multi-vendor, monitoring
Cost optimisation, processing time, result management
Over 1M contacts
Distributed pipeline, real-time verification on ingest, multi-vendor with failover
Infrastructure complexity, vendor management, cost at scale