Data Quality Scoring for Email Lists: Frameworks, Metrics and Implementation
On this page
Why Score Data Quality
Most organisations describe their data quality as "good" or "bad" without measuring it. This makes it impossible to prioritise improvement, track progress or justify investment in data cleaning.
A data quality score assigns a number to each record, segment or list. That number tells you:
- Which records need attention. Sort by score and clean the worst records first.
- Which sources produce good data. Compare scores by acquisition source to identify the best and worst channels.
- Whether quality is improving or declining. Track the average score over time.
- When to stop cleaning. Diminishing returns become visible when the score plateaus.
- What the business impact is. Correlate quality scores with deliverability, engagement and conversion.
Scoring Dimensions
Data quality is not a single attribute. It has multiple dimensions, each of which can be measured and scored independently.
Completeness
Does the record have all the fields you need?
| Field | Weight | Scoring |
|---|---|---|
| Email address | Required (record invalid without it) | Present = 1, missing = 0 (record excluded) |
| First name | High | Present = 1, missing = 0 |
| Last name | High | Present = 1, missing = 0 |
| Company | Medium | Present = 1, missing = 0 |
| Job title | Medium | Present = 1, missing = 0 |
| Phone | Low | Present = 1, missing = 0 |
| Industry | Low | Present = 1, missing = 0 |
| Location | Low | Present = 1, missing = 0 |
Completeness score = sum of (field present x weight) / sum of weights
Example: a record with email, first name, last name and company but missing job title, phone, industry and location:
- Weighted sum: 1(required) + 1(high) + 1(high) + 1(medium) + 0(medium) + 0(low) + 0(low) + 0(low) = meets required, weighted completeness depends on your scheme.
Validity
Is the data in the correct format and does it represent a real value?
| Check | What it validates | Score |
|---|---|---|
| Email syntax | Correct format (user@domain.tld) | Valid = 1, invalid = 0 |
| Email verification | Mailbox exists and accepts email | Verified = 1, unverified = 0.5, invalid = 0 |
| Phone format | Correct number of digits, valid country code | Valid = 1, invalid = 0 |
| Domain exists | MX records present for email domain | Yes = 1, no = 0 |
| Not disposable | Email is not from a disposable domain | Real = 1, disposable = 0 |
| Not role-based | Email is a personal address, not info@ or sales@ | Personal = 1, role-based = 0.5 |
Accuracy
Does the data reflect reality?
This is the hardest dimension to measure because it requires external verification:
- Does the person still work at the listed company?
- Is the job title current?
- Is the phone number still active?
- Is the address correct?
Proxy measures for accuracy:
- Recency of verification. Verified in the last 30 days = 1, 30-90 days = 0.8, 90-180 days = 0.6, over 180 days = 0.4, never = 0.2.
- Engagement history. Opened or clicked in the last 90 days = 1 (proves the address is real and active), 90-180 days = 0.7, over 180 days = 0.4, never engaged = 0.2.
- Bounce history. Never bounced = 1, soft bounced once = 0.7, soft bounced multiple times = 0.3, hard bounced = 0.
Consistency
Is the data formatted consistently across the list?
| Check | Consistent | Inconsistent |
|---|---|---|
| Name capitalisation | "John Smith" | "john smith", "JOHN SMITH", "john Smith" |
| Email case | "john@example.com" | "John@Example.COM" |
| Phone format | "+1 (555) 123-4567" | "5551234567", "555-123-4567", "1-555-123-4567" |
| Company name | "International Business Machines" | "IBM", "I.B.M.", "Intl Business Machines" |
| Country | "United States" | "US", "USA", "U.S.A.", "United States of America" |
| Date format | "2026-10-08" | "10/08/2026", "08-Oct-2026", "October 8, 2026" |
Consistency scoring: compare each field value against the standard format. Matches = 1, does not match = 0.
Uniqueness
Are there duplicate records?
- Unique record = 1.
- Duplicate record = 0.
Upload your lists to Email Extractor for case-insensitive deduplication across all your source files. See How Email Deduplication Works.
Timeliness
How current is the data?
| Age of data | Score |
|---|---|
| Collected or verified in the last 30 days | 1.0 |
| 30-90 days | 0.8 |
| 90-180 days | 0.6 |
| 180-365 days | 0.4 |
| Over 1 year | 0.2 |
| Unknown age | 0.3 |
Building a Scoring Model
Step 1: Define dimensions and weights
Not every dimension matters equally for every use case. Assign weights based on your business:
For cold email outreach:
| Dimension | Weight | Rationale |
|---|---|---|
| Validity (email verification) | 0.35 | Invalid emails bounce and damage sender reputation |
| Accuracy (engagement/recency) | 0.25 | Outdated contacts waste sends |
| Completeness (name, company, title) | 0.20 | Personalisation requires complete data |
| Uniqueness | 0.10 | Duplicates waste volume and look unprofessional |
| Consistency | 0.05 | Important for CRM but less for outreach |
| Timeliness | 0.05 | Captured by accuracy proxy |
For CRM data management:
| Dimension | Weight | Rationale |
|---|---|---|
| Completeness | 0.25 | Full records enable segmentation and reporting |
| Accuracy | 0.25 | Wrong data leads to wrong decisions |
| Uniqueness | 0.20 | Duplicates corrupt pipeline and revenue reporting |
| Consistency | 0.15 | Consistent data enables reliable segmentation |
| Validity | 0.10 | Less critical if data is for internal use |
| Timeliness | 0.05 | Covered by accuracy checks |
Step 2: Score individual records
Calculate a score for each record:
Record score = sum of (dimension score x dimension weight)
Example record:
| Dimension | Raw score | Weight | Weighted score |
|---|---|---|---|
| Validity | 1.0 (verified, valid format) | 0.35 | 0.35 |
| Accuracy | 0.7 (engaged 4 months ago) | 0.25 | 0.175 |
| Completeness | 0.8 (missing phone and industry) | 0.20 | 0.16 |
| Uniqueness | 1.0 (no duplicates) | 0.10 | 0.10 |
| Consistency | 0.6 (name capitalisation issues) | 0.05 | 0.03 |
| Timeliness | 0.8 (collected 2 months ago) | 0.05 | 0.04 |
| Total | 0.855 |
Step 3: Set thresholds
| Score range | Quality tier | Action |
|---|---|---|
| 0.85-1.00 | Excellent | Send, segment, use for high-value campaigns |
| 0.70-0.84 | Good | Send, monitor engagement |
| 0.50-0.69 | Fair | Clean before sending, enrich missing fields |
| 0.30-0.49 | Poor | Do not send, attempt enrichment, reverify |
| 0.00-0.29 | Critical | Suppress, review for removal |
Step 4: Score the list
Aggregate individual scores to get a list-level score:
| Metric | Calculation |
|---|---|
| Average score | Sum of all record scores / number of records |
| Median score | Middle value when sorted |
| Distribution | Percentage of records in each quality tier |
| Sendable percentage | Percentage of records scoring above the sending threshold |
Healthy list benchmarks:
| Metric | Target |
|---|---|
| Average score | Over 0.75 |
| Percentage in "Excellent" or "Good" | Over 70% |
| Percentage in "Poor" or "Critical" | Under 10% |
| Sendable percentage | Over 80% |
Implementation
Spreadsheet approach (small lists)
For lists under 10,000 records, a spreadsheet works:
- Export your list with all fields.
- Add columns for each dimension score.
- Add formulas to calculate dimension scores:
- Completeness:
=COUNTA(B2:H2)/7(count non-empty fields / total fields). - Validity: manual or based on verification results imported from a service.
- Consistency:
=IF(EXACT(B2, PROPER(B2)), 1, 0)(checks name capitalisation).
- Completeness:
- Add a weighted total column:
=SUMPRODUCT(scores, weights). - Add a quality tier column:
=IF(total>=0.85, "Excellent", IF(total>=0.7, "Good", ...)).
Python approach (larger lists)
import pandas as pd
import re
def score_record(row):
scores = {}
# Completeness
fields = ['first_name', 'last_name', 'company', 'title', 'phone', 'industry']
weights = [0.25, 0.25, 0.2, 0.15, 0.1, 0.05]
completeness = sum(
w for f, w in zip(fields, weights)
if pd.notna(row.get(f)) and str(row.get(f)).strip()
)
scores['completeness'] = completeness
# Validity (email)
email = str(row.get('email', ''))
if re.match(r'^[\w.-]+@[\w.-]+\.\w+$', email):
scores['validity'] = 1.0
else:
scores['validity'] = 0.0
# Uniqueness (requires checking against full dataset, simplified here)
scores['uniqueness'] = row.get('is_unique', 1.0)
# Consistency
first = str(row.get('first_name', ''))
last = str(row.get('last_name', ''))
name_consistent = 1.0 if (first == first.title() and last == last.title()) else 0.5
email_consistent = 1.0 if email == email.lower() else 0.5
scores['consistency'] = (name_consistent + email_consistent) / 2
# Timeliness (based on a 'date_added' field)
# Simplified: would calculate days since date_added
scores['timeliness'] = row.get('timeliness_score', 0.5)
# Accuracy (based on engagement data)
scores['accuracy'] = row.get('accuracy_score', 0.5)
# Weighted total
dimension_weights = {
'validity': 0.35,
'accuracy': 0.25,
'completeness': 0.20,
'uniqueness': 0.10,
'consistency': 0.05,
'timeliness': 0.05
}
total = sum(scores[d] * dimension_weights[d] for d in dimension_weights)
return total
# Apply to dataframe
df = pd.read_csv('email_list.csv')
df['quality_score'] = df.apply(score_record, axis=1)
df['quality_tier'] = pd.cut(
df['quality_score'],
bins=[0, 0.3, 0.5, 0.7, 0.85, 1.0],
labels=['Critical', 'Poor', 'Fair', 'Good', 'Excellent']
)
CRM-native scoring
Some CRMs and data quality tools offer built-in scoring:
| Platform | Scoring capability |
|---|---|
| Salesforce | Data quality dashboards (with add-ons like Validity DemandTools) |
| HubSpot | Contact scoring (custom properties and workflows) |
| Marketo | Data quality scoring through smart campaigns |
| Validity DemandTools | Dedicated data quality scoring for Salesforce |
| Openprise | Data orchestration with quality scoring |
| RingLead | Duplicate detection and quality scoring |
Monitoring and Reporting
Dashboard metrics
Track these metrics over time:
| Metric | What it shows |
|---|---|
| Average quality score | Overall list health trend |
| Score distribution | How many records are in each tier |
| Score by source | Which acquisition channels produce the best data |
| Score by segment | Which segments have quality issues |
| Score change over time | Whether quality is improving or declining |
| New record average score | Quality of incoming data |
| Records below threshold | How many records are unsendable |
Alerting
Set up alerts for:
- Average quality score drops below 0.70.
- More than 15% of records fall into "Poor" or "Critical."
- A new data source produces records with an average score below 0.60.
- The percentage of sendable records drops below 75%.
- A specific segment's quality drops significantly.
Reporting cadence
| Report | Frequency | Audience |
|---|---|---|
| Quality dashboard | Real-time / daily | Marketing operations |
| Source quality report | Weekly | Marketing operations, demand gen |
| Segment quality report | Monthly | Marketing, sales |
| Executive summary | Quarterly | Leadership |
| Annual audit | Yearly | Marketing, sales, compliance |
Continuous Improvement
Using scores to prioritise cleaning
- Sort records by quality score (ascending).
- Identify the most common issues among low-scoring records.
- Address the highest-impact issue first (the one that affects the most records).
- Re-score after cleaning to measure improvement.
- Track the cost per point of quality improvement (time and money spent vs score increase).
Preventing quality decay
Quality decays naturally. People change jobs, companies, email addresses and phone numbers. Prevention is cheaper than cleaning:
- Validate data at the point of entry (form validation, API verification).
- Verify the full list quarterly.
- Remove hard bounces immediately.
- Re-engage or suppress inactive contacts.
- Enrich records periodically (update company, title, phone).
- Score new records on import and reject imports below a threshold.
Benchmarking
Compare your scores to industry benchmarks:
| Industry | Typical average score | Notes |
|---|---|---|
| SaaS/Technology | 0.75-0.85 | Digital-native, frequent engagement |
| Professional services | 0.70-0.80 | Relationship-driven, less frequent data updates |
| E-commerce | 0.65-0.75 | High volume, more disposable addresses |
| Healthcare | 0.60-0.70 | Complex compliance, slower data updates |
| Manufacturing | 0.55-0.70 | Less digital engagement, slower data decay |
These are illustrative ranges based on the scoring framework above, not from a specific published study. Your results will depend on your sources, industry and cleaning practices.