Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Data Quality Scoring for Email Lists: Frameworks, Metrics and Implementation

On this page

Why Score Data Quality

Most organisations describe their data quality as "good" or "bad" without measuring it. This makes it impossible to prioritise improvement, track progress or justify investment in data cleaning.

A data quality score assigns a number to each record, segment or list. That number tells you:

  • Which records need attention. Sort by score and clean the worst records first.
  • Which sources produce good data. Compare scores by acquisition source to identify the best and worst channels.
  • Whether quality is improving or declining. Track the average score over time.
  • When to stop cleaning. Diminishing returns become visible when the score plateaus.
  • What the business impact is. Correlate quality scores with deliverability, engagement and conversion.

Scoring Dimensions

Data quality is not a single attribute. It has multiple dimensions, each of which can be measured and scored independently.

Completeness

Does the record have all the fields you need?

Field Weight Scoring
Email address Required (record invalid without it) Present = 1, missing = 0 (record excluded)
First name High Present = 1, missing = 0
Last name High Present = 1, missing = 0
Company Medium Present = 1, missing = 0
Job title Medium Present = 1, missing = 0
Phone Low Present = 1, missing = 0
Industry Low Present = 1, missing = 0
Location Low Present = 1, missing = 0

Completeness score = sum of (field present x weight) / sum of weights

Example: a record with email, first name, last name and company but missing job title, phone, industry and location:

  • Weighted sum: 1(required) + 1(high) + 1(high) + 1(medium) + 0(medium) + 0(low) + 0(low) + 0(low) = meets required, weighted completeness depends on your scheme.

Validity

Is the data in the correct format and does it represent a real value?

Check What it validates Score
Email syntax Correct format (user@domain.tld) Valid = 1, invalid = 0
Email verification Mailbox exists and accepts email Verified = 1, unverified = 0.5, invalid = 0
Phone format Correct number of digits, valid country code Valid = 1, invalid = 0
Domain exists MX records present for email domain Yes = 1, no = 0
Not disposable Email is not from a disposable domain Real = 1, disposable = 0
Not role-based Email is a personal address, not info@ or sales@ Personal = 1, role-based = 0.5

Accuracy

Does the data reflect reality?

This is the hardest dimension to measure because it requires external verification:

  • Does the person still work at the listed company?
  • Is the job title current?
  • Is the phone number still active?
  • Is the address correct?

Proxy measures for accuracy:

  • Recency of verification. Verified in the last 30 days = 1, 30-90 days = 0.8, 90-180 days = 0.6, over 180 days = 0.4, never = 0.2.
  • Engagement history. Opened or clicked in the last 90 days = 1 (proves the address is real and active), 90-180 days = 0.7, over 180 days = 0.4, never engaged = 0.2.
  • Bounce history. Never bounced = 1, soft bounced once = 0.7, soft bounced multiple times = 0.3, hard bounced = 0.

Consistency

Is the data formatted consistently across the list?

Check Consistent Inconsistent
Name capitalisation "John Smith" "john smith", "JOHN SMITH", "john Smith"
Email case "john@example.com" "John@Example.COM"
Phone format "+1 (555) 123-4567" "5551234567", "555-123-4567", "1-555-123-4567"
Company name "International Business Machines" "IBM", "I.B.M.", "Intl Business Machines"
Country "United States" "US", "USA", "U.S.A.", "United States of America"
Date format "2026-10-08" "10/08/2026", "08-Oct-2026", "October 8, 2026"

Consistency scoring: compare each field value against the standard format. Matches = 1, does not match = 0.

Uniqueness

Are there duplicate records?

  • Unique record = 1.
  • Duplicate record = 0.

Upload your lists to Email Extractor for case-insensitive deduplication across all your source files. See How Email Deduplication Works.

Timeliness

How current is the data?

Age of data Score
Collected or verified in the last 30 days 1.0
30-90 days 0.8
90-180 days 0.6
180-365 days 0.4
Over 1 year 0.2
Unknown age 0.3

Building a Scoring Model

Step 1: Define dimensions and weights

Not every dimension matters equally for every use case. Assign weights based on your business:

For cold email outreach:

Dimension Weight Rationale
Validity (email verification) 0.35 Invalid emails bounce and damage sender reputation
Accuracy (engagement/recency) 0.25 Outdated contacts waste sends
Completeness (name, company, title) 0.20 Personalisation requires complete data
Uniqueness 0.10 Duplicates waste volume and look unprofessional
Consistency 0.05 Important for CRM but less for outreach
Timeliness 0.05 Captured by accuracy proxy

For CRM data management:

Dimension Weight Rationale
Completeness 0.25 Full records enable segmentation and reporting
Accuracy 0.25 Wrong data leads to wrong decisions
Uniqueness 0.20 Duplicates corrupt pipeline and revenue reporting
Consistency 0.15 Consistent data enables reliable segmentation
Validity 0.10 Less critical if data is for internal use
Timeliness 0.05 Covered by accuracy checks

Step 2: Score individual records

Calculate a score for each record:

Record score = sum of (dimension score x dimension weight)

Example record:

Dimension Raw score Weight Weighted score
Validity 1.0 (verified, valid format) 0.35 0.35
Accuracy 0.7 (engaged 4 months ago) 0.25 0.175
Completeness 0.8 (missing phone and industry) 0.20 0.16
Uniqueness 1.0 (no duplicates) 0.10 0.10
Consistency 0.6 (name capitalisation issues) 0.05 0.03
Timeliness 0.8 (collected 2 months ago) 0.05 0.04
Total 0.855

Step 3: Set thresholds

Score range Quality tier Action
0.85-1.00 Excellent Send, segment, use for high-value campaigns
0.70-0.84 Good Send, monitor engagement
0.50-0.69 Fair Clean before sending, enrich missing fields
0.30-0.49 Poor Do not send, attempt enrichment, reverify
0.00-0.29 Critical Suppress, review for removal

Step 4: Score the list

Aggregate individual scores to get a list-level score:

Metric Calculation
Average score Sum of all record scores / number of records
Median score Middle value when sorted
Distribution Percentage of records in each quality tier
Sendable percentage Percentage of records scoring above the sending threshold

Healthy list benchmarks:

Metric Target
Average score Over 0.75
Percentage in "Excellent" or "Good" Over 70%
Percentage in "Poor" or "Critical" Under 10%
Sendable percentage Over 80%

Implementation

Spreadsheet approach (small lists)

For lists under 10,000 records, a spreadsheet works:

  1. Export your list with all fields.
  2. Add columns for each dimension score.
  3. Add formulas to calculate dimension scores:
    • Completeness: =COUNTA(B2:H2)/7 (count non-empty fields / total fields).
    • Validity: manual or based on verification results imported from a service.
    • Consistency: =IF(EXACT(B2, PROPER(B2)), 1, 0) (checks name capitalisation).
  4. Add a weighted total column: =SUMPRODUCT(scores, weights).
  5. Add a quality tier column: =IF(total>=0.85, "Excellent", IF(total>=0.7, "Good", ...)).

Python approach (larger lists)

import pandas as pd
import re

def score_record(row):
    scores = {}

    # Completeness
    fields = ['first_name', 'last_name', 'company', 'title', 'phone', 'industry']
    weights = [0.25, 0.25, 0.2, 0.15, 0.1, 0.05]
    completeness = sum(
        w for f, w in zip(fields, weights)
        if pd.notna(row.get(f)) and str(row.get(f)).strip()
    )
    scores['completeness'] = completeness

    # Validity (email)
    email = str(row.get('email', ''))
    if re.match(r'^[\w.-]+@[\w.-]+\.\w+$', email):
        scores['validity'] = 1.0
    else:
        scores['validity'] = 0.0

    # Uniqueness (requires checking against full dataset, simplified here)
    scores['uniqueness'] = row.get('is_unique', 1.0)

    # Consistency
    first = str(row.get('first_name', ''))
    last = str(row.get('last_name', ''))
    name_consistent = 1.0 if (first == first.title() and last == last.title()) else 0.5
    email_consistent = 1.0 if email == email.lower() else 0.5
    scores['consistency'] = (name_consistent + email_consistent) / 2

    # Timeliness (based on a 'date_added' field)
    # Simplified: would calculate days since date_added
    scores['timeliness'] = row.get('timeliness_score', 0.5)

    # Accuracy (based on engagement data)
    scores['accuracy'] = row.get('accuracy_score', 0.5)

    # Weighted total
    dimension_weights = {
        'validity': 0.35,
        'accuracy': 0.25,
        'completeness': 0.20,
        'uniqueness': 0.10,
        'consistency': 0.05,
        'timeliness': 0.05
    }

    total = sum(scores[d] * dimension_weights[d] for d in dimension_weights)
    return total

# Apply to dataframe
df = pd.read_csv('email_list.csv')
df['quality_score'] = df.apply(score_record, axis=1)
df['quality_tier'] = pd.cut(
    df['quality_score'],
    bins=[0, 0.3, 0.5, 0.7, 0.85, 1.0],
    labels=['Critical', 'Poor', 'Fair', 'Good', 'Excellent']
)

CRM-native scoring

Some CRMs and data quality tools offer built-in scoring:

Platform Scoring capability
Salesforce Data quality dashboards (with add-ons like Validity DemandTools)
HubSpot Contact scoring (custom properties and workflows)
Marketo Data quality scoring through smart campaigns
Validity DemandTools Dedicated data quality scoring for Salesforce
Openprise Data orchestration with quality scoring
RingLead Duplicate detection and quality scoring

Monitoring and Reporting

Dashboard metrics

Track these metrics over time:

Metric What it shows
Average quality score Overall list health trend
Score distribution How many records are in each tier
Score by source Which acquisition channels produce the best data
Score by segment Which segments have quality issues
Score change over time Whether quality is improving or declining
New record average score Quality of incoming data
Records below threshold How many records are unsendable

Alerting

Set up alerts for:

  • Average quality score drops below 0.70.
  • More than 15% of records fall into "Poor" or "Critical."
  • A new data source produces records with an average score below 0.60.
  • The percentage of sendable records drops below 75%.
  • A specific segment's quality drops significantly.

Reporting cadence

Report Frequency Audience
Quality dashboard Real-time / daily Marketing operations
Source quality report Weekly Marketing operations, demand gen
Segment quality report Monthly Marketing, sales
Executive summary Quarterly Leadership
Annual audit Yearly Marketing, sales, compliance

Continuous Improvement

Using scores to prioritise cleaning

  1. Sort records by quality score (ascending).
  2. Identify the most common issues among low-scoring records.
  3. Address the highest-impact issue first (the one that affects the most records).
  4. Re-score after cleaning to measure improvement.
  5. Track the cost per point of quality improvement (time and money spent vs score increase).

Preventing quality decay

Quality decays naturally. People change jobs, companies, email addresses and phone numbers. Prevention is cheaper than cleaning:

  • Validate data at the point of entry (form validation, API verification).
  • Verify the full list quarterly.
  • Remove hard bounces immediately.
  • Re-engage or suppress inactive contacts.
  • Enrich records periodically (update company, title, phone).
  • Score new records on import and reject imports below a threshold.

Benchmarking

Compare your scores to industry benchmarks:

Industry Typical average score Notes
SaaS/Technology 0.75-0.85 Digital-native, frequent engagement
Professional services 0.70-0.80 Relationship-driven, less frequent data updates
E-commerce 0.65-0.75 High volume, more disposable addresses
Healthcare 0.60-0.70 Complex compliance, slower data updates
Manufacturing 0.55-0.70 Less digital engagement, slower data decay

These are illustrative ranges based on the scoring framework above, not from a specific published study. Your results will depend on your sources, industry and cleaning practices.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)