Data Anonymisation and Pseudonymisation for Email Datasets
By Email ExtractorPublished 9 min read
On this page
Why Anonymise Email Data
Email datasets frequently need to be shared, analysed or stored in ways that require removing or obscuring personal information:
Scenario
Why anonymisation is needed
Sharing data with vendors
Vendor does not need to see actual email addresses to do their work
Testing and development
Developers and QA need realistic data without real personal information
Analytics and reporting
Aggregate analysis does not require individual identification
Regulatory compliance (GDPR, CCPA)
Data minimisation principle requires limiting personal data exposure
Academic research
Research datasets should not contain identifiable personal information
Data breach preparation
Anonymised data that is breached has lower impact
Cross-border data transfer
Anonymised data may not be subject to data transfer restrictions
Archival
Long-term storage of anonymised data avoids ongoing compliance obligations
Anonymisation vs Pseudonymisation
These are different techniques with different legal implications:
Aspect
Anonymisation
Pseudonymisation
Definition
Irreversibly removing all identifying information
Replacing identifiers with artificial ones; reversal is possible with a key
Reversible
No
Yes (with the key)
GDPR status
Not personal data (GDPR does not apply)
Still personal data (GDPR applies)
Data utility
Lower (less analytical value)
Higher (structure preserved; can be re-identified for updates)
Risk
Re-identification risk must be negligible
Key must be secured; risk if key is compromised
Use case
Public datasets, irreversible sharing
Internal analysis, testing, vendor sharing with contractual protections
GDPR definitions
Term
GDPR definition
Implication for email data
Personal data
Any information relating to an identified or identifiable person
Email addresses are personal data
Anonymisation
Processing so that data can no longer be attributed to a specific person
Truly anonymised data falls outside GDPR scope
Pseudonymisation
Processing so that data can no longer be attributed without additional information
Pseudonymised data is still personal data under GDPR
Anonymisation Techniques for Email Data
Technique comparison
Technique
How it works
Reversible
Data utility
Re-identification risk
Deletion
Remove email addresses entirely
No
None for email analysis
None
Hashing
Replace email with a hash (SHA-256, etc.)
Not directly, but vulnerable to lookup attacks
Can link records by same hash
Medium (lookup tables)
Salted hashing
Hash with a secret salt
No (without the salt)
Can link records by same hash
Low (if salt is secure)
Tokenisation
Replace email with a random token; store mapping separately
Yes (with mapping table)
Can link records by same token
Low (if mapping is secure)
Generalisation
Replace email with domain only (keep @company.com, remove local part)
No
Company-level analysis only
Depends on domain uniqueness
Synthetic replacement
Replace with realistic but fake email addresses
No
Structure preserved for testing
None (data is fictitious)
Masking
Replace characters (j***@example.com)
No
Partial recognition; limited analysis
Low
K-anonymity
Ensure each record is indistinguishable from at least k-1 others
No
Reduced but usable
Controlled by k value
When to use each technique
Scenario
Recommended technique
Rationale
Public dataset release
Deletion or synthetic replacement
Strongest protection; no re-identification risk
Internal analytics
Salted hashing or tokenisation
Preserves ability to link records; controlled access
Development and testing
Synthetic replacement
Realistic data without any real personal information
Vendor data sharing
Tokenisation
Vendor can work with data; mapping stays with you
Cross-border transfer
Anonymisation (deletion or synthetic)
Removes GDPR data transfer restrictions
Aggregate reporting
Generalisation (domain only)
Company-level insights without individual identification
Backup and archival
Pseudonymisation (tokenisation)
Can restore if needed; reduced exposure if breached
Implementation
Hashing email addresses
import hashlib
import secrets
def hash_email(email, salt=None):
"""Hash an email address. Use a salt for security."""
email = email.strip().lower()
if salt:
data = (salt + email).encode('utf-8')
else:
data = email.encode('utf-8')
return hashlib.sha256(data).hexdigest()
# Without salt (vulnerable to lookup attacks)
hashed = hash_email('user@example.com')
# Result: a consistent hash; same input always produces same output
# With salt (more secure)
salt = secrets.token_hex(16) # Generate once; store securely
hashed = hash_email('user@example.com', salt=salt)
# Result: different hash than unsalted; cannot be reversed without salt
Warning about unsalted hashing: Unsalted hashes of email addresses are vulnerable to lookup attacks. An attacker can hash a list of known email addresses and compare them to your hashed dataset. Always use a salt, and keep the salt secret.
Tokenisation
import uuid
import csv
class EmailTokeniser:
"""Replace email addresses with random tokens."""
def __init__(self):
self.mapping = {} # email -> token
self.reverse_mapping = {} # token -> email
def tokenise(self, email):
"""Replace an email with a consistent random token."""
email = email.strip().lower()
if email not in self.mapping:
token = str(uuid.uuid4())
self.mapping[email] = token
self.reverse_mapping[token] = email
return self.mapping[email]
def detokenise(self, token):
"""Recover the original email from a token."""
return self.reverse_mapping.get(token)
def process_csv(self, input_file, output_file, email_column='email'):
"""Tokenise email addresses in a CSV file."""
with open(input_file, 'r') as inf, \
open(output_file, 'w', newline='') as outf:
reader = csv.DictReader(inf)
writer = csv.DictWriter(outf, fieldnames=reader.fieldnames)
writer.writeheader()
for row in reader:
if email_column in row and row[email_column]:
row[email_column] = self.tokenise(
row[email_column])
writer.writerow(row)
def save_mapping(self, mapping_file):
"""Save the mapping table (store securely)."""
with open(mapping_file, 'w', newline='') as f:
writer = csv.writer(f)
writer.writerow(['token', 'email'])
for token, email in self.reverse_mapping.items():
writer.writerow([token, email])
Synthetic email generation
import random
import string
def generate_synthetic_email(domain='example.com'):
"""Generate a realistic but fake email address."""
first_names = [
'alex', 'jordan', 'taylor', 'morgan', 'casey',
'riley', 'avery', 'quinn', 'blake', 'drew'
]
last_names = [
'smith', 'jones', 'wilson', 'brown', 'taylor',
'davis', 'miller', 'anderson', 'thomas', 'jackson'
]
separators = ['.', '_', '']
first = random.choice(first_names)
last = random.choice(last_names)
sep = random.choice(separators)
number = random.randint(1, 999) if random.random() > 0.5 else ''
return f"{first}{sep}{last}{number}@{domain}"
def replace_emails_with_synthetic(emails, preserve_domains=False):
"""Replace real emails with synthetic ones."""
synthetic = {}
for email in emails:
email_lower = email.strip().lower()
if email_lower not in synthetic:
if preserve_domains:
domain = email_lower.split('@')[1]
else:
domain = 'example.com'
synthetic[email_lower] = generate_synthetic_email(domain)
return synthetic
Domain-only generalisation
def generalise_to_domain(email):
"""Remove the local part; keep only the domain."""
parts = email.strip().lower().split('@')
if len(parts) == 2:
return f"[redacted]@{parts[1]}"
return "[invalid]"
# "john.smith@acmecorp.com" becomes "[redacted]@acmecorp.com"
Re-Identification Risks
Even anonymised data can sometimes be re-identified:
Risk
How it works
Mitigation
Unsalted hash lookup
Attacker hashes common email addresses and compares to your dataset
A domain with one employee makes "[redacted]@smallcompany.com" identifiable
Suppress domains below a threshold size
Linkage attack
Linking your anonymised dataset with another dataset that has identifying information
Assess linkage risks before sharing
Inference attack
Deducing identity from patterns in the data
Review data for unique patterns
Temporal correlation
Timestamps correlate with known activities
Generalise timestamps (date only, not time)
Assessing re-identification risk
Factor
Lower risk
Higher risk
Dataset size
Large (millions of records)
Small (hundreds of records)
Field diversity
Many possible values per field
Few possible values (easy to narrow down)
External data availability
No public datasets to link against
Many public datasets available for linking
Domain diversity
Many different domains in dataset
Few domains (especially small organisations)
Temporal precision
Generalised dates (month/year)
Exact timestamps
Geographic precision
Country or region
Exact location
Workflow for Anonymising Email Datasets
Step-by-step process
Step
Action
Purpose
1. Inventory
List all fields containing personal data
Know what needs to be processed
2. Classify
Categorise each field (direct identifier, quasi-identifier, non-identifying)
Determine treatment for each field
3. Choose technique
Select the appropriate technique for each field
Balance utility and privacy
4. Process
Apply the chosen technique to each field
Create the anonymised dataset
5. Validate
Check that anonymisation is complete and correct
Ensure no personal data remains
6. Assess risk
Evaluate re-identification risk
Confirm the level of anonymisation is sufficient
7. Document
Record what was done and why
Compliance and audit trail
8. Secure
Store any mapping tables or salts with appropriate security
Prevent re-identification if pseudonymised
Using Email Extractor in the workflow
Before anonymising, you may need to extract email addresses from unstructured data:
Upload source files (CSV, XLSX, PDF, TXT, etc.) to Email Extractor to extract all email addresses.
The tool's automatic deduplication (case-insensitive) produces a clean starting list.
Download as CSV, noting which source each address came from.
Apply the chosen anonymisation technique to the extracted list.
Use the anonymised version for analysis, sharing or testing.
Note that Email Extractor processes data client-side in the browser, meaning the extracted email addresses are not sent to a server. This is relevant for compliance: the extraction step itself does not create an additional data processing activity involving a third party.
Compliance Considerations
Requirement
How anonymisation helps
GDPR data minimisation
Anonymised datasets contain only necessary information
GDPR right to erasure
Truly anonymised data does not need to be erased (it is no longer personal data)
GDPR data transfer
Truly anonymised data is not subject to cross-border transfer restrictions
CCPA right to delete
Anonymised data may fall outside CCPA scope
Data retention policies
Anonymised data can be retained longer than personal data
Breach notification
Breach of anonymised data may not trigger notification requirements
Vendor data sharing
Anonymised data can be shared with fewer contractual restrictions
Pseudonymisation under GDPR
Requirement
Implementation
Separate storage
Store the mapping table separately from the pseudonymised data
Access control
Restrict access to the mapping table
Technical measures
Encrypt the mapping table at rest and in transit
Organisational measures
Document who has access to the mapping table and why
Purpose limitation
Use re-identification only for the stated purpose
Data subject rights
Pseudonymised data is still personal data; rights still apply