Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Automating Data Validation for Email Lists: Rules, Tools and Pipelines

On this page

Why Automate Validation

Manual data validation does not scale. A marketing team adding 500 contacts per week cannot manually check each one for formatting errors, duplicates, disposable domains and invalid syntax. Errors slip through, list quality degrades and deliverability suffers.

Automated validation catches problems at the point of entry, before bad data contaminates the rest of the system. It runs the same checks every time, does not get tired and does not skip steps when the team is busy.

What to Validate

Syntax validation

Syntax validation checks whether an email address is correctly formed according to the email specification (RFC 5321/5322).

Rules:

  • Contains exactly one @ symbol.
  • Local part (before @) is 1-64 characters.
  • Domain part (after @) is 1-255 characters.
  • Domain contains at least one dot (with exceptions for internal domains).
  • No spaces or prohibited characters.
  • Does not start or end with a dot or hyphen in the domain.

Common syntax errors caught:

  • Missing @ symbol (user.example.com).
  • Double @ symbols (user@@example.com).
  • Spaces in the address (user @example.com).
  • Missing domain extension (user@example).
  • Special characters in wrong positions (user@.example.com).

Regex pattern (covers most cases):

^[a-zA-Z0-9.!#$%&'*+/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*\.[a-zA-Z]{2,}$

See Email Syntax Validation and Regex for Email Extraction.

Domain validation

Domain validation checks whether the domain in an email address can receive mail.

Rules:

  • Domain has DNS records (the domain exists).
  • Domain has MX records (mail servers are configured).
  • Domain is not on a known blocklist.
  • Domain is not a known disposable email provider.
  • Domain is not a typo of a common domain (gmial.com, yaho.com, outlok.com).

Role-based address detection

Role-based addresses (info@, sales@, support@, admin@, webmaster@) are not personal addresses. They typically go to a shared inbox or distribution list.

Why flag them:

  • Higher bounce rates (some are unmonitored).
  • Higher spam complaint rates (multiple people see the message).
  • Lower engagement rates.
  • Some ESPs flag or reject them.

Common role-based prefixes: info, sales, support, admin, webmaster, postmaster, abuse, noreply, no-reply, marketing, hr, billing, contact, help, enquiries, office, team, hello, careers, press, media, legal, compliance.

See Role-Based Addresses.

Disposable address detection

Disposable email services (Guerrilla Mail, Temp Mail, 10MinuteMail, Mailinator and hundreds of others) provide temporary addresses that expire after minutes or hours.

Detection methods:

  • Domain blocklist (maintained lists of known disposable domains).
  • Pattern matching (many disposable services use recognisable domain patterns).
  • DNS analysis (some disposable services share infrastructure).

See Disposable Addresses.

Duplicate detection

Duplicates waste send volume and create confusion in reporting.

Deduplication rules:

Upload your lists to Email Extractor for case-insensitive deduplication across multiple files and formats.

See How Email Deduplication Works and Removing Duplicates at Scale.

Format normalisation

Normalise data at entry to prevent downstream inconsistencies.

Email normalisation:

  • Lowercase the entire address.
  • Trim leading and trailing whitespace.
  • Remove invisible Unicode characters.
  • Standardise encoding (convert punycode to Unicode or vice versa as needed).

Associated field normalisation:

  • Names: title case, trim whitespace, handle prefixes/suffixes.
  • Phone: consistent format (E.164 for international).
  • Company: standardise common abbreviations (Inc., LLC, Ltd.).
  • Country: ISO 3166-1 alpha-2 codes.

See Normalizing Email Formats and Data Normalization Workflows.

Suppression list matching

Every new contact should be checked against your suppression list before being added to any sendable list.

Suppression list sources:

  • Previous unsubscribes.
  • Hard bounces.
  • Spam complaints.
  • Manual removal requests (data subject requests).
  • Legal do-not-contact lists.
  • Company-wide suppression (competitors, current customers to be excluded from prospecting).

Automation Architecture

Pipeline design

A validation pipeline processes each contact through a series of checks in a defined order. The order matters because early checks can prevent unnecessary (and sometimes costly) later checks.

Recommended order:

  1. Format normalisation. Lowercase, trim whitespace, remove invisible characters. No cost, instant.
  2. Syntax validation. Regex check. No cost, instant. Reject or flag invalid syntax.
  3. Duplicate check. Compare against existing database. No cost beyond database query.
  4. Suppression list check. Compare against suppression list. No cost beyond database query.
  5. Domain validation. DNS/MX lookup. Minimal cost, sub-second. Reject if domain does not exist.
  6. Disposable domain check. Compare against blocklist. No cost, instant.
  7. Role-based check. Compare prefix against known role-based prefixes. No cost, instant.
  8. Typo detection. Compare domain against common typos. No cost, instant. Suggest correction.
  9. Email verification. SMTP check or verification API. Cost per check (typically $0.003-$0.01). Takes 1-5 seconds.

Why this order: Steps 1-8 are free and fast. They eliminate obviously bad data before step 9, which costs money and takes time. If 20% of incoming contacts fail steps 1-8, you save 20% on verification costs.

Entry point validation

Validate at every point where contacts enter the system:

Web forms:

  • Client-side: syntax validation, typo suggestions (instant feedback).
  • Server-side: all checks through step 8 (before saving to database).
  • Async: email verification (step 9) after submission, before adding to sendable list.

Bulk imports (CSV, spreadsheet uploads):

  • Pre-import: all checks (1-9) on the entire file.
  • Report: summary of results (valid, invalid, duplicates, disposable, role-based).
  • Action: import only valid contacts; quarantine the rest for review.

API integrations (CRM sync, form tools, third-party data):

  • Middleware: validation service between the source and your database.
  • All checks (1-9) before the contact reaches the database.
  • Logging: record every contact received, every check performed, every decision made.

Manual entry:

  • Form validation on the data entry screen.
  • Same rules as web forms.

Implementation options

Option 1: Built-in ESP/CRM validation

Most ESPs and CRMs include basic validation:

Platform Built-in validation
Mailchimp Syntax, some duplicate checking
HubSpot Syntax, duplicate detection, some domain checking
Salesforce Duplicate rules, validation rules (configurable)
ActiveCampaign Syntax, bounce processing

Limitation: Built-in validation is usually limited to syntax and basic duplicate checking. It does not include disposable domain detection, role-based checking, typo detection or SMTP verification.

Option 2: Validation API services

Dedicated email validation APIs provide comprehensive checking:

Service Checks included Pricing model
ZeroBounce Syntax, domain, SMTP, disposable, role-based, abuse, catch-all Per verification
NeverBounce Syntax, domain, SMTP, disposable, accept-all Per verification
Bouncer Syntax, domain, SMTP, disposable, role-based, toxicity Per verification
Kickbox Syntax, domain, SMTP, disposable, role-based, accept-all Per verification
Hunter.io Syntax, domain, SMTP, disposable, webmail Per verification

Integration approach:

import requests

def validate_email(email):
    # Step 1-8: Local checks (free, instant)
    if not passes_syntax_check(email):
        return {'valid': False, 'reason': 'invalid_syntax'}

    if is_duplicate(email):
        return {'valid': False, 'reason': 'duplicate'}

    if is_suppressed(email):
        return {'valid': False, 'reason': 'suppressed'}

    if is_disposable_domain(email):
        return {'valid': False, 'reason': 'disposable'}

    if is_role_based(email):
        return {'valid': True, 'warning': 'role_based'}

    typo = detect_typo(email)
    if typo:
        return {'valid': True, 'warning': 'possible_typo', 'suggestion': typo}

    # Step 9: API verification (paid, slower)
    response = requests.get(
        'https://api.example.com/verify',
        params={'email': email, 'api_key': 'YOUR_KEY'}
    )
    result = response.json()

    return {
        'valid': result['status'] == 'valid',
        'reason': result.get('sub_status', ''),
        'risk': result.get('risk_score', 0)
    }

See Email Validation API Integration and Email Verification Service Comparison.

Option 3: Custom validation pipeline

For teams with development resources, build a custom pipeline using open-source components:

Components:

  • Syntax validation: regex or a library (Python: email-validator; JavaScript: validator.js).
  • Domain validation: DNS library for MX lookups (Python: dnspython; JavaScript: dns module).
  • Disposable domain list: Open-source lists (disposable-email-domains on GitHub, updated regularly).
  • Role-based list: Maintain your own list of role-based prefixes.
  • Typo detection: Levenshtein distance against common domains.
  • Deduplication: Database-level unique constraints or application-level sets.
  • SMTP verification: Build or use a library (caution: many mail servers block SMTP probes).

Option 4: Integration platform automation

No-code platforms can orchestrate validation workflows:

Zapier workflow example:

  1. Trigger: new form submission (Typeform, Google Forms, etc.).
  2. Action: send email to validation API (ZeroBounce, NeverBounce).
  3. Filter: if valid, continue; if invalid, send to quarantine spreadsheet.
  4. Action: check against suppression list (Google Sheets lookup or CRM query).
  5. Action: add valid, non-suppressed contact to CRM/ESP.

Make (Integromat) workflow example:

  1. Trigger: new row in Google Sheets or new webhook.
  2. HTTP module: call validation API.
  3. Router: branch based on validation result.
  4. Valid path: create contact in CRM.
  5. Invalid path: log to error spreadsheet with reason.

See Zapier Email Extraction Workflows and Make Scenarios Email Extraction Workflows.

Error Handling

What to do with invalid contacts

Not all invalid contacts should be discarded. Handle based on the type of failure:

Failure type Action
Invalid syntax Reject. Attempt auto-correction (missing .com, obvious typo). If corrected, re-validate.
Domain does not exist Reject. No recovery possible.
MX records missing Quarantine. Domain may be misconfigured temporarily. Re-check in 48 hours.
Disposable domain Reject for marketing lists. May be acceptable for one-time transactional purposes.
Role-based Accept with flag. Some teams exclude from cold outreach but include in customer communication.
Duplicate Merge. Update existing record with any new information from the duplicate.
Suppressed Reject. Do not add to any sendable list.
SMTP undeliverable Reject. Move to bounce list.
SMTP catch-all Accept with flag. The domain accepts all addresses, so validity cannot be confirmed.
Possible typo Accept with flag. Present suggestion to user if interactive; log for review if batch.

Quarantine workflow

Contacts that fail validation but might be recoverable go to quarantine:

  1. Store the contact with the failure reason and timestamp.
  2. For recoverable failures (temporary DNS issues, SMTP timeouts), re-validate after 48 hours.
  3. For ambiguous failures (catch-all domains, possible typos), flag for manual review.
  4. Set a quarantine expiry (30 days). Contacts not resolved within that period are deleted.
  5. Report quarantine volume and resolution rates as part of data quality metrics.

Monitoring and Reporting

Key metrics

Metric Target Action if exceeded
Invalid rate (new contacts) Under 5% Review acquisition sources with high invalid rates
Duplicate rate (new contacts) Under 10% Review data entry processes and integrations
Disposable rate (new contacts) Under 2% Add disposable detection to forms
Role-based rate (new contacts) Under 15% Depends on acquisition source (directories may be higher)
Suppression match rate Under 1% Review how suppressed contacts are re-entering
Validation API error rate Under 1% Check API service status, review error logs

Alerting

Set up alerts for:

  • Validation API failure (service down or returning errors).
  • Unusual spike in invalid contacts (may indicate a bot attack on forms).
  • Suppression matches above threshold (may indicate a data flow problem).
  • Quarantine growing beyond expected size.
  • Validation processing backlog (batch jobs taking longer than expected).

Dashboards

Build a validation dashboard showing:

  • Contacts processed today/this week/this month.
  • Pass/fail breakdown by validation step.
  • Top failure reasons.
  • Validation results by source (which forms or integrations produce the most invalid data).
  • Cost tracking (validation API spend).

Implementation Steps

Week 1: Audit

  1. List every point where contacts enter your system (forms, imports, integrations, manual entry).
  2. For each entry point, document what validation currently exists (if any).
  3. Pull a sample of 1,000 recent contacts and run them through a validation service to establish a baseline invalid rate.
  4. Upload the sample to Email Extractor to check for duplicates in the sample.

Week 2: Design

  1. Define validation rules for each check (syntax, domain, disposable, role-based, duplicate, suppression).
  2. Choose the implementation approach (built-in, API, custom or integration platform).
  3. Design the pipeline order.
  4. Define error handling rules (reject, quarantine, accept with flag).
  5. Define monitoring and alerting requirements.

Week 3-4: Build

  1. Implement local checks first (syntax, duplicate, suppression, disposable, role-based). These are free and provide immediate value.
  2. Integrate a validation API for SMTP verification.
  3. Build or configure error handling and quarantine workflows.
  4. Set up monitoring and alerting.
  5. Test with a batch of known good and known bad addresses.

Week 5: Deploy

  1. Enable validation on one entry point (start with the highest-volume source).
  2. Monitor results for a week. Adjust rules if the false-positive rate is too high.
  3. Roll out to remaining entry points.
  4. Set up regular reporting.

Common Mistakes

Validating only at import, not at entry. If contacts are added to the CRM without validation and only validated during periodic list cleaning, bad data is already in the system and may have been used.

Validating once and never again. Email addresses go bad over time. People leave companies, domains expire, mailboxes are deactivated. Re-validate your full list quarterly.

Treating all failures the same. A role-based address is not the same as a syntax error. A catch-all domain is not the same as a non-existent domain. Nuanced handling produces better results than blanket rejection.

Not monitoring validation costs. Validation API charges are per-check. Without monitoring, costs can spike unexpectedly, especially if a form is hit by bots submitting thousands of fake addresses.

Blocking legitimate contacts. Over-aggressive validation (rejecting all free email providers, blocking all role-based addresses) can reject legitimate prospects. Calibrate rules to your specific use case.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)