Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Data Anonymisation and Pseudonymisation for Email Datasets

On this page

Why Anonymise Email Data

Email datasets frequently need to be shared, analysed or stored in ways that require removing or obscuring personal information:

Scenario Why anonymisation is needed
Sharing data with vendors Vendor does not need to see actual email addresses to do their work
Testing and development Developers and QA need realistic data without real personal information
Analytics and reporting Aggregate analysis does not require individual identification
Regulatory compliance (GDPR, CCPA) Data minimisation principle requires limiting personal data exposure
Academic research Research datasets should not contain identifiable personal information
Data breach preparation Anonymised data that is breached has lower impact
Cross-border data transfer Anonymised data may not be subject to data transfer restrictions
Archival Long-term storage of anonymised data avoids ongoing compliance obligations

Anonymisation vs Pseudonymisation

These are different techniques with different legal implications:

Aspect Anonymisation Pseudonymisation
Definition Irreversibly removing all identifying information Replacing identifiers with artificial ones; reversal is possible with a key
Reversible No Yes (with the key)
GDPR status Not personal data (GDPR does not apply) Still personal data (GDPR applies)
Data utility Lower (less analytical value) Higher (structure preserved; can be re-identified for updates)
Risk Re-identification risk must be negligible Key must be secured; risk if key is compromised
Use case Public datasets, irreversible sharing Internal analysis, testing, vendor sharing with contractual protections

GDPR definitions

Term GDPR definition Implication for email data
Personal data Any information relating to an identified or identifiable person Email addresses are personal data
Anonymisation Processing so that data can no longer be attributed to a specific person Truly anonymised data falls outside GDPR scope
Pseudonymisation Processing so that data can no longer be attributed without additional information Pseudonymised data is still personal data under GDPR

Anonymisation Techniques for Email Data

Technique comparison

Technique How it works Reversible Data utility Re-identification risk
Deletion Remove email addresses entirely No None for email analysis None
Hashing Replace email with a hash (SHA-256, etc.) Not directly, but vulnerable to lookup attacks Can link records by same hash Medium (lookup tables)
Salted hashing Hash with a secret salt No (without the salt) Can link records by same hash Low (if salt is secure)
Tokenisation Replace email with a random token; store mapping separately Yes (with mapping table) Can link records by same token Low (if mapping is secure)
Generalisation Replace email with domain only (keep @company.com, remove local part) No Company-level analysis only Depends on domain uniqueness
Synthetic replacement Replace with realistic but fake email addresses No Structure preserved for testing None (data is fictitious)
Masking Replace characters (j***@example.com) No Partial recognition; limited analysis Low
K-anonymity Ensure each record is indistinguishable from at least k-1 others No Reduced but usable Controlled by k value

When to use each technique

Scenario Recommended technique Rationale
Public dataset release Deletion or synthetic replacement Strongest protection; no re-identification risk
Internal analytics Salted hashing or tokenisation Preserves ability to link records; controlled access
Development and testing Synthetic replacement Realistic data without any real personal information
Vendor data sharing Tokenisation Vendor can work with data; mapping stays with you
Cross-border transfer Anonymisation (deletion or synthetic) Removes GDPR data transfer restrictions
Aggregate reporting Generalisation (domain only) Company-level insights without individual identification
Backup and archival Pseudonymisation (tokenisation) Can restore if needed; reduced exposure if breached

Implementation

Hashing email addresses

import hashlib
import secrets

def hash_email(email, salt=None):
    """Hash an email address. Use a salt for security."""
    email = email.strip().lower()
    if salt:
        data = (salt + email).encode('utf-8')
    else:
        data = email.encode('utf-8')
    return hashlib.sha256(data).hexdigest()

# Without salt (vulnerable to lookup attacks)
hashed = hash_email('user@example.com')
# Result: a consistent hash; same input always produces same output

# With salt (more secure)
salt = secrets.token_hex(16)  # Generate once; store securely
hashed = hash_email('user@example.com', salt=salt)
# Result: different hash than unsalted; cannot be reversed without salt

Warning about unsalted hashing: Unsalted hashes of email addresses are vulnerable to lookup attacks. An attacker can hash a list of known email addresses and compare them to your hashed dataset. Always use a salt, and keep the salt secret.

Tokenisation

import uuid
import csv

class EmailTokeniser:
    """Replace email addresses with random tokens."""

    def __init__(self):
        self.mapping = {}  # email -> token
        self.reverse_mapping = {}  # token -> email

    def tokenise(self, email):
        """Replace an email with a consistent random token."""
        email = email.strip().lower()
        if email not in self.mapping:
            token = str(uuid.uuid4())
            self.mapping[email] = token
            self.reverse_mapping[token] = email
        return self.mapping[email]

    def detokenise(self, token):
        """Recover the original email from a token."""
        return self.reverse_mapping.get(token)

    def process_csv(self, input_file, output_file, email_column='email'):
        """Tokenise email addresses in a CSV file."""
        with open(input_file, 'r') as inf, \
             open(output_file, 'w', newline='') as outf:
            reader = csv.DictReader(inf)
            writer = csv.DictWriter(outf, fieldnames=reader.fieldnames)
            writer.writeheader()
            for row in reader:
                if email_column in row and row[email_column]:
                    row[email_column] = self.tokenise(
                        row[email_column])
                writer.writerow(row)

    def save_mapping(self, mapping_file):
        """Save the mapping table (store securely)."""
        with open(mapping_file, 'w', newline='') as f:
            writer = csv.writer(f)
            writer.writerow(['token', 'email'])
            for token, email in self.reverse_mapping.items():
                writer.writerow([token, email])

Synthetic email generation

import random
import string

def generate_synthetic_email(domain='example.com'):
    """Generate a realistic but fake email address."""
    first_names = [
        'alex', 'jordan', 'taylor', 'morgan', 'casey',
        'riley', 'avery', 'quinn', 'blake', 'drew'
    ]
    last_names = [
        'smith', 'jones', 'wilson', 'brown', 'taylor',
        'davis', 'miller', 'anderson', 'thomas', 'jackson'
    ]
    separators = ['.', '_', '']

    first = random.choice(first_names)
    last = random.choice(last_names)
    sep = random.choice(separators)
    number = random.randint(1, 999) if random.random() > 0.5 else ''

    return f"{first}{sep}{last}{number}@{domain}"

def replace_emails_with_synthetic(emails, preserve_domains=False):
    """Replace real emails with synthetic ones."""
    synthetic = {}
    for email in emails:
        email_lower = email.strip().lower()
        if email_lower not in synthetic:
            if preserve_domains:
                domain = email_lower.split('@')[1]
            else:
                domain = 'example.com'
            synthetic[email_lower] = generate_synthetic_email(domain)
    return synthetic

Domain-only generalisation

def generalise_to_domain(email):
    """Remove the local part; keep only the domain."""
    parts = email.strip().lower().split('@')
    if len(parts) == 2:
        return f"[redacted]@{parts[1]}"
    return "[invalid]"

# "john.smith@acmecorp.com" becomes "[redacted]@acmecorp.com"

Re-Identification Risks

Even anonymised data can sometimes be re-identified:

Risk How it works Mitigation
Unsalted hash lookup Attacker hashes common email addresses and compares to your dataset Use salted hashing; keep salt secret
Quasi-identifier combination Combining anonymised fields (domain, timestamp, location) uniquely identifies someone Remove or generalise quasi-identifiers
Small domain size A domain with one employee makes "[redacted]@smallcompany.com" identifiable Suppress domains below a threshold size
Linkage attack Linking your anonymised dataset with another dataset that has identifying information Assess linkage risks before sharing
Inference attack Deducing identity from patterns in the data Review data for unique patterns
Temporal correlation Timestamps correlate with known activities Generalise timestamps (date only, not time)

Assessing re-identification risk

Factor Lower risk Higher risk
Dataset size Large (millions of records) Small (hundreds of records)
Field diversity Many possible values per field Few possible values (easy to narrow down)
External data availability No public datasets to link against Many public datasets available for linking
Domain diversity Many different domains in dataset Few domains (especially small organisations)
Temporal precision Generalised dates (month/year) Exact timestamps
Geographic precision Country or region Exact location

Workflow for Anonymising Email Datasets

Step-by-step process

Step Action Purpose
1. Inventory List all fields containing personal data Know what needs to be processed
2. Classify Categorise each field (direct identifier, quasi-identifier, non-identifying) Determine treatment for each field
3. Choose technique Select the appropriate technique for each field Balance utility and privacy
4. Process Apply the chosen technique to each field Create the anonymised dataset
5. Validate Check that anonymisation is complete and correct Ensure no personal data remains
6. Assess risk Evaluate re-identification risk Confirm the level of anonymisation is sufficient
7. Document Record what was done and why Compliance and audit trail
8. Secure Store any mapping tables or salts with appropriate security Prevent re-identification if pseudonymised

Using Email Extractor in the workflow

Before anonymising, you may need to extract email addresses from unstructured data:

  1. Upload source files (CSV, XLSX, PDF, TXT, etc.) to Email Extractor to extract all email addresses.
  2. The tool's automatic deduplication (case-insensitive) produces a clean starting list.
  3. Download as CSV, noting which source each address came from.
  4. Apply the chosen anonymisation technique to the extracted list.
  5. Use the anonymised version for analysis, sharing or testing.

Note that Email Extractor processes data client-side in the browser, meaning the extracted email addresses are not sent to a server. This is relevant for compliance: the extraction step itself does not create an additional data processing activity involving a third party.

Compliance Considerations

Requirement How anonymisation helps
GDPR data minimisation Anonymised datasets contain only necessary information
GDPR right to erasure Truly anonymised data does not need to be erased (it is no longer personal data)
GDPR data transfer Truly anonymised data is not subject to cross-border transfer restrictions
CCPA right to delete Anonymised data may fall outside CCPA scope
Data retention policies Anonymised data can be retained longer than personal data
Breach notification Breach of anonymised data may not trigger notification requirements
Vendor data sharing Anonymised data can be shared with fewer contractual restrictions

Pseudonymisation under GDPR

Requirement Implementation
Separate storage Store the mapping table separately from the pseudonymised data
Access control Restrict access to the mapping table
Technical measures Encrypt the mapping table at rest and in transit
Organisational measures Document who has access to the mapping table and why
Purpose limitation Use re-identification only for the stated purpose
Data subject rights Pseudonymised data is still personal data; rights still apply

Testing with Anonymised Data

Creating test datasets

Approach Method Pros Cons
Anonymised production data Anonymise a copy of real data Realistic data distribution Must ensure complete anonymisation
Synthetic data Generate entirely fake data No privacy risk May not reflect real-world patterns
Hybrid Anonymised structure with synthetic values Realistic structure; no real data More complex to create
Subset and anonymise Take a small sample and anonymise Manageable size; realistic Sample may not be representative

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)