Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Building an Email Validation Pipeline: From Raw Data to Clean Lists

On this page

What Is a Validation Pipeline?

A validation pipeline is a series of automated steps that take raw email data and produce a clean, verified, ready-to-use list. Instead of manually uploading CSVs to different tools and downloading results, the pipeline runs end to end with minimal human intervention.

Why build one:

  • Consistency. Every email that enters your system goes through the same checks.
  • Speed. New data is cleaned within hours, not days.
  • Scale. The pipeline handles 1,000 or 1,000,000 addresses the same way.
  • Quality. Catch problems before they reach your CRM, email platform or sales team.

Pipeline Stages

Stage 1: Ingestion

Collect email data from all sources into a single staging area.

Sources to connect:

  • Form submissions (website signups, lead magnets, event registrations).
  • File uploads (CSV, XLSX, PDF imports from partners, data providers, purchased lists).
  • CRM exports (periodic full exports or incremental syncs).
  • API feeds (data provider APIs, marketing platform APIs).
  • Manual entry (sales rep input, business card data).

Implementation:

  • For file-based sources, extract email addresses using Email Extractor or a programmatic extraction step (regex or a parsing library for structured formats).
  • For API sources, build connectors that pull new contacts on a schedule.
  • For form submissions, capture directly from the form processor webhook.
  • Write all ingested records to a staging table or file with: email, source, timestamp and any associated metadata.

Stage 2: Normalisation

Standardise email formatting before validation.

Steps:

  1. Trim whitespace. Remove leading, trailing and internal spaces.
  2. Lowercase. Convert the entire address to lowercase.
  3. Strip plus-addressing. Remove +tag portions (user+tag@example.com becomes user@example.com). Optional, depending on your use case.
  4. Fix common typos. Correct known domain misspellings:
    • gnail.com to gmail.com
    • yaho.com to yahoo.com
    • outlok.com to outlook.com
    • hotmal.com to hotmail.com
  5. Remove invalid characters. Strip characters that are not allowed in standard email addresses.
  6. Handle encoding. Convert HTML entities, URL encoding and unicode normalisation.
import re

def normalise_email(email):
    if not email:
        return None

    # Trim and lowercase
    email = email.strip().lower()

    # Remove spaces within the address
    email = email.replace(' ', '')

    # Fix common domain typos
    domain_fixes = {
        'gnail.com': 'gmail.com',
        'gmial.com': 'gmail.com',
        'gmal.com': 'gmail.com',
        'gamil.com': 'gmail.com',
        'yaho.com': 'yahoo.com',
        'yahooo.com': 'yahoo.com',
        'hotmal.com': 'hotmail.com',
        'hotmial.com': 'hotmail.com',
        'outlok.com': 'outlook.com',
        'outllook.com': 'outlook.com',
    }

    parts = email.split('@')
    if len(parts) != 2:
        return None

    local, domain = parts
    domain = domain_fixes.get(domain, domain)

    return f"{local}@{domain}"

See Normalizing Email Formats.

Stage 3: Deduplication

Remove duplicate entries before spending money on verification.

Deduplication levels:

  1. Exact match. Same normalised address appears multiple times.
  2. Cross-source. Same address from multiple sources (keep the richest record).
  3. Gmail dot variants. For @gmail.com addresses, j.doe and jdoe deliver to the same mailbox.

Implementation:

  • Group by normalised email address.
  • When duplicates exist, keep the record with the most complete metadata, or the most recent.
  • Track which sources contributed each address (useful for source quality analysis).

See How Email Deduplication Works.

Stage 4: Syntax validation

Check whether the email address follows valid syntax before making any network calls. This is free and instant.

Checks:

  • Contains exactly one @ sign.
  • Local part is not empty.
  • Domain part is not empty.
  • Domain contains at least one dot.
  • TLD is at least 2 characters.
  • No consecutive dots in the domain.
  • No prohibited characters.
import re

def is_valid_syntax(email):
    pattern = r'^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$'
    return bool(re.match(pattern, email))

Action: Remove addresses that fail syntax validation. They will always bounce.

See Email Syntax Validation.

Stage 5: DNS/MX validation

Check whether the domain has mail servers configured to receive email.

How it works:

  1. Extract the domain from the email address.
  2. Query DNS for MX (Mail Exchange) records.
  3. If MX records exist, the domain can receive email.
  4. If no MX records, check for an A record (some domains accept email without MX records).
  5. If neither exists, the domain cannot receive email.
import dns.resolver

def has_mx_records(domain):
    try:
        records = dns.resolver.resolve(domain, 'MX')
        return len(records) > 0
    except (dns.resolver.NoAnswer,
            dns.resolver.NXDOMAIN,
            dns.resolver.NoNameservers,
            dns.exception.Timeout):
        return False

Action: Remove addresses with no MX or A records. The domain does not accept email.

Performance tip: Cache MX lookup results by domain. Multiple addresses at the same domain need only one DNS query.

See MX Records Explained.

Stage 6: Disposable and role-based detection

Identify addresses that are valid but low-quality for outreach.

Disposable addresses: Temporary email services (Guerrilla Mail, Temp Mail, Mailinator). Addresses expire within hours or days.

Role-based addresses: Shared inboxes (info@, admin@, sales@, support@, webmaster@). Not tied to an individual.

Implementation: Maintain or subscribe to lists of known disposable email domains and role-based local parts.

DISPOSABLE_DOMAINS = {'guerrillamail.com', 'tempmail.com',
                      'mailinator.com', 'throwaway.email', ...}

ROLE_PREFIXES = {'info', 'admin', 'sales', 'support', 'help',
                 'contact', 'webmaster', 'postmaster', 'abuse',
                 'noreply', 'no-reply', 'billing', 'marketing'}

def classify_email(email):
    local, domain = email.split('@')
    flags = []

    if domain in DISPOSABLE_DOMAINS:
        flags.append('disposable')

    if local in ROLE_PREFIXES:
        flags.append('role_based')

    return flags

Action: Flag (do not necessarily remove) disposable and role-based addresses. The decision to include or exclude depends on the use case.

See Disposable Email Addresses and Role-Based Addresses.

Stage 7: SMTP verification

Connect to the recipient's mail server and check whether the specific mailbox exists. This is the most definitive check but also the slowest and most resource-intensive.

How it works:

  1. Connect to the MX server on port 25.
  2. Send HELO/EHLO command.
  3. Send MAIL FROM command.
  4. Send RCPT TO command with the email address.
  5. The server responds with acceptance (250) or rejection (550).
  6. Disconnect without sending a message.

Limitations:

  • Catch-all domains. Some servers accept all addresses regardless. SMTP verification returns "valid" even for non-existent mailboxes.
  • Greylisting. Some servers temporarily reject the first connection attempt.
  • Rate limiting. Servers may block or delay repeated verification attempts.
  • Privacy-conscious servers. Some servers do not reveal whether a mailbox exists (they accept all RCPT TO commands).

Action: For most teams, using a third-party verification API (which handles SMTP verification at scale, manages IP reputation and handles retries) is more practical than building SMTP verification in-house.

See SMTP Email Verification Explained.

Stage 8: Third-party verification API

For production pipelines, send addresses through a verification API that handles all of the above checks plus additional signals.

Integration:

def verify_batch(emails, api_key):
    """Send batch to verification service."""
    response = requests.post(
        'https://api.verificationservice.example.com/v1/batch',
        json={'emails': emails},
        headers={'Authorization': f'Bearer {api_key}'},
        timeout=30
    )
    return response.json()['batch_id']

# Submit, wait, retrieve
batch_id = verify_batch(email_list, API_KEY)
results = poll_for_results(batch_id)

See Email Validation API Integration and Email Verification Service Comparison.

Stage 9: Scoring and categorisation

Assign a quality score to each address based on the validation results.

def score_email(result):
    score = 0

    if result['valid']:
        score += 50
    if not result['disposable']:
        score += 15
    if not result['role_based']:
        score += 10
    if not result['free_provider']:
        score += 10
    if not result['catch_all']:
        score += 15

    return score

Categorisation:

Score Category Action
80-100 High quality Add to active outreach lists
50-79 Medium quality Add with caution, monitor bounce rates
20-49 Low quality Use for low-risk campaigns only
0-19 Poor quality Suppress or remove

Stage 10: Output and routing

Route validated addresses to the appropriate destination:

  • High-quality addresses go to the CRM and outreach sequences.
  • Medium-quality addresses go to nurture lists.
  • Invalid addresses go to a rejection log for analysis.
  • Disposable and role-based addresses go to a flagged list for manual review.

Automation and Scheduling

Real-time pipeline

For form submissions and API integrations, run validation in real time:

  1. Form submits email address.
  2. Pipeline runs stages 2-8 in under 3 seconds.
  3. Valid address is accepted; invalid is rejected with a user-friendly error.

Batch pipeline

For file uploads and periodic cleaning, run on a schedule:

  1. New files land in an ingestion folder.
  2. A scheduled job runs the pipeline every hour (or daily).
  3. Results are delivered to the output destination.

Periodic re-verification

Schedule re-verification of your entire database every 3-6 months:

  1. Export all active contacts.
  2. Run through stages 5-9.
  3. Flag newly invalid addresses for removal.
  4. Report on list decay rate.

See Email List Decay Explained.

Monitoring

Track pipeline health with these metrics:

  • Volume. How many addresses processed per day/week.
  • Pass rate. What percentage of addresses pass validation.
  • Source quality. Pass rate broken down by source (which sources produce the cleanest data).
  • Processing time. How long the pipeline takes (important for real-time use cases).
  • API costs. Verification API spend per period.
  • Bounce rate (post-pipeline). Addresses that passed validation but still bounced when emailed. This is your pipeline's false positive rate.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)