Data Quality Frameworks: Building a Systematic Approach to Contact Data Quality
On this page
What Is a Data Quality Framework?
A data quality framework is a structured approach to defining, measuring, monitoring and improving the quality of data across an organisation. Rather than treating data cleaning as a one-off project, a framework turns it into a repeatable process with clear ownership, standards and metrics.
For contact data -- email addresses, phone numbers, company information, job titles -- quality directly affects deliverability, personalisation, segmentation and revenue attribution. A single duplicate or invalid email address costs little on its own, but thousands of them degrade every system they touch.
The Six Dimensions of Data Quality
Most frameworks measure quality across six dimensions. Each applies differently to contact and email data:
Accuracy
Does the data correctly represent the real-world entity?
| Contact data example | Accurate | Inaccurate |
|---|---|---|
| Email address | Current, valid, reaches the intended person | Misspelled, belongs to a different person, no longer active |
| Job title | Current role at current company | Previous role, wrong company |
| Company name | Legal or commonly used name | Abbreviation that matches a different company |
| Phone number | Current, reaches the intended person | Old number, wrong area code |
How to measure: Compare a random sample of records against a verified source (LinkedIn profile, company website, direct confirmation). Calculate the percentage that match.
Target: 90%+ accuracy for actively used contact records. 95%+ for records entering automated outreach.
Completeness
Are all required fields populated?
Define required fields by use case:
| Use case | Required fields |
|---|---|
| Email outreach | Email, first name, company |
| Account-based marketing | Email, first name, last name, company, job title, industry |
| Sales prospecting | Email, first name, last name, company, job title, phone, LinkedIn URL |
| Event follow-up | Email, first name, event attended, session attended |
| Customer support | Email, account ID, company |
How to measure: Percentage of records with all required fields populated (not just non-null -- a field containing "N/A" or "unknown" is not complete).
Target: 95%+ completeness for required fields in active records.
Consistency
Is the same entity represented the same way across systems?
| Inconsistency | System A | System B |
|---|---|---|
| Name format | "Smith, John" | "John Smith" |
| Company name | "IBM" | "International Business Machines" |
| Country | "US" | "United States" |
| Email case | "John@Example.com" | "john@example.com" |
| Date format | "10/08/2026" | "2026-10-08" |
How to measure: Select records that exist in multiple systems. Calculate the percentage where key fields match exactly after normalisation.
Target: 100% consistency for identifiers (email, account ID). 90%+ for descriptive fields.
Timeliness
Is the data current enough for its intended use?
| Data type | Decay rate | Refresh frequency needed |
|---|---|---|
| Email address | 2-3% per month become invalid | Verify quarterly |
| Job title | 20-30% change annually | Refresh every 6 months |
| Phone number | 10-15% change annually | Refresh annually |
| Company information | Mergers, acquisitions, closures | Monitor continuously |
| Mailing address | 10-15% change annually | Refresh annually |
How to measure: Track the age of each record (days since last verified or updated). Calculate the percentage of records within their acceptable age window.
Target: 80%+ of active records verified within the last 6 months.
Uniqueness
Is each real-world entity represented exactly once?
Duplicate detection for contact data is harder than it appears:
| Duplicate type | Example | Detection method |
|---|---|---|
| Exact email duplicate | Same email in two records | Exact match after lowercasing |
| Near-duplicate email | john+newsletter@example.com and john@example.com | Plus-tag stripping, Gmail dot handling |
| Same person, different email | john@example.com (personal) and j.smith@example.org (work) | Name + company matching |
| Same company, different spelling | "Acme Corp" and "ACME Corporation" | Fuzzy matching, domain matching |
How to measure: Run deduplication logic across the full database. Calculate duplicates as a percentage of total records.
Target: Under 2% duplicate rate for the primary contact database.
Upload contact exports from multiple systems to Email Extractor to identify duplicate email addresses across sources. The tool's case-insensitive deduplication catches exact duplicates automatically. Download results as CSV with sources to inspect the source names for each address. Source overlap concerns exact email strings; matching people or resolving conflicting CRM fields requires the original contact records.
Validity
Does the data conform to defined formats and business rules?
| Field | Validity rule |
|---|---|
| Contains exactly one @, valid domain with MX record, local part under 64 characters | |
| Phone | Matches E.164 format, valid country code |
| Country | ISO 3166-1 alpha-2 code |
| Date | ISO 8601 format (YYYY-MM-DD) |
| URL | Valid scheme (https://), resolvable domain |
| Job title | Not a placeholder ("test", "asdf", "N/A") |
How to measure: Run validation rules against all records. Calculate the percentage that pass.
Target: 99%+ validity for structured fields (email, phone, country). 95%+ for semi-structured fields (job title, company name).
Assessment Methodology
Step 1: Define scope
Decide what data to assess:
| Scope | What to assess |
|---|---|
| Single system | One CRM, one marketing platform, one database |
| Cross-system | All systems that store contact data |
| Single segment | One list, one campaign audience, one territory |
| Full database | Every contact record across all systems |
Start with a single system or segment. A full cross-system assessment is valuable but significantly more complex.
Step 2: Profile the data
Before measuring quality, understand what you have:
Record counts:
- Total records.
- Records by source (imported, form submission, purchased, scraped, manual entry).
- Records by age (created date).
- Records by last activity (last email sent, last email opened, last form submission).
Field analysis:
- Population rate for each field (percentage non-null).
- Value distribution for each field (top values, outliers).
- Format consistency (how many phone numbers match a standard format?).
Relationship analysis:
- Records per company (are there 500 contacts at a 10-person company?).
- Contacts per deal/opportunity.
- Cross-system overlap (how many CRM contacts also exist in the marketing platform?).
Step 3: Measure against dimensions
For each dimension, run the measurement described above. Record results in a scorecard:
| Dimension | Metric | Score | Target | Gap |
|---|---|---|---|---|
| Accuracy | % verified correct (sample) | 82% | 90% | -8% |
| Completeness | % with all required fields | 71% | 95% | -24% |
| Consistency | % matching across systems | 64% | 90% | -26% |
| Timeliness | % verified in last 6 months | 45% | 80% | -35% |
| Uniqueness | Duplicate rate | 8% | 2% | +6% |
| Validity | % passing format rules | 93% | 99% | -6% |
Step 4: Prioritise improvements
Not all dimensions matter equally for every use case. Prioritise by business impact:
| If your problem is... | Prioritise... |
|---|---|
| High bounce rates | Accuracy, validity, timeliness |
| Low personalisation rates | Completeness |
| Duplicate outreach (same person contacted multiple times) | Uniqueness, consistency |
| Conflicting reports across teams | Consistency |
| Wasted budget on bad records | All dimensions |
Governance Structure
Roles and responsibilities
| Role | Responsibility | Typical owner |
|---|---|---|
| Data owner | Defines quality standards for their domain | VP Sales, VP Marketing, CRO |
| Data steward | Implements and monitors quality rules | Marketing Ops, Sales Ops, RevOps |
| Data custodian | Manages technical infrastructure | IT, Data Engineering |
| Data consumer | Reports quality issues, follows data entry standards | Sales reps, marketers, SDRs |
| Executive sponsor | Funds and prioritises data quality initiatives | CRO, COO, CFO |
Data quality policies
Document policies for:
Data entry standards:
- Required fields by record type (lead, contact, account).
- Naming conventions (title case for names, specific format for company names).
- Validation rules enforced at point of entry.
- Duplicate prevention rules (what to do when a duplicate is detected during entry).
Data maintenance:
- Scheduled verification cycles (quarterly email verification, semi-annual enrichment).
- Bounce handling (what happens when an email bounces -- immediate suppression? Re-verification?).
- Role-based address handling (info@, sales@, support@ -- include or exclude?).
- Record archival rules (when does a record move from active to archived?).
Data acquisition:
- Approved sources (which list vendors, enrichment providers, scraping tools are approved?).
- Minimum quality requirements for imported data.
- Consent and compliance requirements (GDPR, CAN-SPAM, CCPA).
- Onboarding process for new data sources.
RACI matrix for common data operations
| Operation | Responsible | Accountable | Consulted | Informed |
|---|---|---|---|---|
| Import new list | Marketing Ops | Data owner | Legal/Compliance | Sales |
| Merge duplicates | RevOps | Data steward | Sales reps (for owned accounts) | Marketing |
| Bulk update job titles | RevOps | Data steward | Sales | Marketing |
| Add new data source | Data steward | Data owner | IT, Legal | All consumers |
| Archive inactive records | RevOps | Data owner | Sales, Marketing | Finance |
| Respond to data subject request | Legal/Compliance | DPO | IT, Marketing | Data owner |
Continuous Improvement
Monitoring dashboard
Track these metrics on a regular cadence:
Weekly:
- New records added (by source).
- Bounce rate from email sends.
- Duplicate records created.
- Validation failures at point of entry.
Monthly:
- Overall quality score (composite of six dimensions).
- Quality score trend (improving, stable, declining).
- Records by completeness tier (fully complete, partially complete, minimal).
- Duplicate rate.
Quarterly:
- Full quality assessment against targets.
- Source quality comparison (which sources produce the highest and lowest quality records?).
- Cost of poor data quality (estimated wasted spend on bad records).
- Progress on improvement initiatives.
Root cause analysis
When quality scores decline, investigate the source:
| Symptom | Possible root causes |
|---|---|
| Rising bounce rate | Stale list, bad data source, email verification lapsed |
| Increasing duplicates | New integration without dedup rules, manual entry without duplicate check |
| Declining completeness | New required fields not enforced, bulk imports without all fields |
| Cross-system inconsistency | Integration sync failure, manual edits in one system not propagated |
| Invalid format entries | Form validation disabled, API integration sending malformed data |
Improvement cycle
Use a Plan-Do-Check-Act (PDCA) cycle:
- Plan. Identify the highest-impact quality issue from the latest assessment. Define the improvement goal, approach and timeline.
- Do. Implement the fix. This might be a one-time cleanup (deduplicate the database), a process change (add validation to a form), or a tool implementation (add email verification to the import workflow).
- Check. Measure the dimension again after the fix. Did it move the score toward the target?
- Act. If the fix worked, standardise it (make it part of the ongoing process). If not, analyse why and adjust.
Tools for Data Quality
Assessment and profiling
| Tool | What it does | Best for |
|---|---|---|
| Great Expectations (open source) | Define, test and document data quality expectations | Engineering teams with Python pipelines |
| Ataccama | Data profiling, cataloguing, quality rules | Enterprise data governance |
| Informatica Data Quality | Profiling, standardisation, matching | Enterprise data management |
| Talend Data Quality | Profiling, cleansing, matching | Mid-market, Talend ecosystem |
| Monte Carlo | Data observability, anomaly detection | Detecting quality regressions in pipelines |
| dbt tests | SQL-based data quality assertions | Teams already using dbt for transformation |
Email-specific tools
| Tool | Function |
|---|---|
| Email Extractor | Extract and deduplicate emails from files and text |
| ZeroBounce, NeverBounce, Bouncer | Email verification (mailbox-level validation) |
| Clearbit, ZoomInfo, Apollo | Contact enrichment (add missing fields) |
| Clay | Waterfall enrichment (try multiple sources) |
| Normative | Email format standardisation |
CRM-native tools
| CRM | Built-in quality features |
|---|---|
| Salesforce | Duplicate management rules, validation rules, data quality dashboards (CRM Analytics) |
| HubSpot | Duplicate management, property validation, data quality command centre |
| Zoho CRM | Deduplication, validation rules, data enrichment (Zia) |
| Dynamics 365 | Duplicate detection rules, data quality dashboards |
Implementation Roadmap
Phase 1: Foundation (weeks 1-4)
- Appoint a data steward.
- Define required fields by record type.
- Run initial data quality assessment.
- Produce a scorecard with baseline measurements.
- Identify the top 3 quality issues by business impact.
Phase 2: Quick wins (weeks 5-8)
- Deduplicate the primary contact database.
- Run email verification on the active list.
- Add validation rules to CRM fields and web forms.
- Fix the top data entry issues (training, templates, defaults).
- Set up weekly bounce rate monitoring.
Phase 3: Process (weeks 9-16)
- Document data entry standards and distribute to the team.
- Set up automated quality checks on data imports.
- Implement cross-system consistency rules (integration-level standardisation).
- Create a monthly quality dashboard.
- Establish a regular enrichment and verification cycle.
Phase 4: Governance (weeks 17-24)
- Formalise data quality policies.
- Assign data owners for each data domain.
- Implement the RACI matrix for data operations.
- Set up quarterly quality reviews with stakeholders.
- Establish a feedback loop for data consumers to report issues.
Phase 5: Optimisation (ongoing)
- Automate quality scoring and alerting.
- Implement real-time quality checks (validate at point of entry, not after the fact).
- Track cost of poor data quality and ROI of quality improvements.
- Expand quality framework to new data domains as the organisation grows.
Common Pitfalls
Treating data quality as a one-off project. A single cleanup decays within months. Build ongoing processes, not projects.
No executive sponsor. Without leadership buy-in, data quality competes (and loses) against revenue-generating priorities.
Measuring everything, fixing nothing. A dashboard that shows declining quality without triggering action is worse than no dashboard -- it creates the illusion of governance.
Perfectionism. 100% accuracy across every field is not achievable or necessary. Focus on the fields and records that drive the most business value.
Ignoring data entry. Most quality problems originate at the point of entry. Fixing downstream symptoms without addressing the source is a losing strategy.
Siloed ownership. Sales owns CRM data, marketing owns the email platform, support owns the ticket system. Without cross-functional governance, each team optimises its own system while cross-system quality degrades.