Using Email Extraction During Data Migration
On this page
When Migration Requires Extraction
Platform migrations rarely go smoothly. When you move from one CRM, email service provider or internal system to another, the clean CSV export you expected often turns out to be a collection of inconsistent files in different formats. You may end up with:
- Database exports in CSV or TSV with inconsistent column names across tables.
- Old spreadsheets in XLSX, XLS or ODS that were never imported into the original system.
- Archived emails in EML or MSG format containing contact information in signatures and message bodies.
- PDF reports, invoices or contracts with email addresses embedded in the text.
- JSON or XML data exports from APIs or legacy systems.
- VCF files exported from address books.
- Text logs from old help desk or ticketing systems.
When the goal is to consolidate every email address from these scattered sources into a single, clean list for import into a new platform, extracting addresses from each file type and deduplicating the results is the most direct path.
Common Migration Scenarios
CRM to CRM
Moving from one CRM to another often involves exporting contacts, deals, notes and activity logs. The primary contact export usually provides a clean list, but email addresses also appear in:
- Activity logs and call notes where a contact's colleague was mentioned by email.
- Deal records where a secondary contact was added in a text field rather than as a linked contact.
- Attached documents like proposals, contracts or order confirmations.
Extracting from these supplementary files catches addresses that exist in the old CRM's data but were never formally added as contacts.
Email service provider migration
When moving from one email marketing platform to another, you typically export subscriber lists as CSV files. But addresses also live in:
- Bounce logs and suppression lists (often exported as CSV or TXT).
- Campaign reports in PDF or spreadsheet format.
- Automated email templates saved as HTML files.
Extracting from all of these ensures you migrate not just active subscribers but also your suppression data, which is important for compliance and deliverability.
Legacy system retirement
Older systems may not have a structured export option at all. You may be working with:
- Database dumps in text or log format.
- Saved emails from the system's notification output (EML or MSG files).
- Printed reports that were later scanned (requiring OCR before extraction).
- XML configuration files that contain administrator or user email addresses.
Step-by-Step Migration Extraction
1. Gather all source files
Before extracting, collect every file from the old system that could contain email addresses. Cast a wide net. It is easier to deduplicate a comprehensive list than to discover missing addresses after migration.
Common export locations:
- The platform's built-in export or backup feature.
- Database export tools (producing CSV, JSON, XML or SQL dumps).
- File storage or document management systems.
- Email archives.
- Local backups on team members' machines.
2. Organise files by type
Sort your collected files by format. This helps you understand what you are working with and identifies any files that need preprocessing:
- Directly supported formats (upload to Email Extractor as-is): TXT, TSV, CSV, HTML, HTM, XML, JSON, MD, LOG, VCF, DOCX, XLSX, XLSM, XLSB, XLS, ODS, PDF, MSG, EML.
- Formats requiring conversion first: MBOX files need to be converted to individual EML files. ZIP archives need to be unzipped before selecting supported files.
- Formats requiring external tools first: Scanned PDFs need OCR processing before extraction. Database binary files need to be exported to a supported text format.
3. Extract in batches
Email Extractor accepts files up to 25 MB each and 100 MB per batch. For a large migration, work through your files in batches grouped by source or type.
To extract:
- Go to Email Extractor.
- Select "Text and files."
- Upload a batch of files.
- Click "Extract emails."
Download results as CSV with sources for each batch. The source column records which file each address was found in, which is valuable during migration for tracking provenance.
4. Consolidate and deduplicate
After extracting from all batches, combine your CSV files. You will likely have duplicate addresses across batches, especially if the same contact appeared in multiple systems or file types.
You can paste the combined text back into Email Extractor to deduplicate. The tool removes duplicates using case-insensitive matching, so User@Example.com and user@example.com are treated as the same address.
5. Clean the consolidated list
Before importing into your new platform:
- Remove role-based addresses like info@, support@ and noreply@ unless you specifically need them. See Role-based email addresses.
- Check for disposable addresses that will not be useful in the new system. See Disposable email addresses.
- Cross-reference with your suppression list. Addresses that unsubscribed or bounced in the old system should not be treated as active subscribers in the new one. See What is an email suppression list.
- Validate addresses through an email validation service to catch invalid or expired addresses before importing. Email Extractor does not validate addresses.
6. Map to your new platform's format
Each CRM or email platform expects data in a specific format. Once you have a clean, deduplicated list, format it to match your new platform's import requirements. See How to format your email list for a CRM.
Preserving Context During Migration
Email Extractor outputs email addresses. It does not preserve names, company fields, tags or other metadata from the source files. If your migration requires associated data (names, companies, deal stages), you will need to handle that mapping separately.
However, the CSV with sources download helps bridge this gap. By recording which file each address came from, you can trace an address back to its source document and manually associate context where needed.
For example, if you extract from a file called acme-corp-contract-2025.pdf and see procurement@example.com in the results with that file as its source, you know the context for that address.
Handling Large Migrations
For migrations involving hundreds or thousands of files:
Prioritise by file type. Start with structured data (CSV, XLSX, VCF) because these typically contain the most reliable address data. Move to semi-structured formats (JSON, XML) next, then unstructured text (PDF, DOCX, TXT, LOG) last.
Keep a log. Track which files you have processed, how many addresses each batch produced and any files that need preprocessing. This prevents duplicate work and helps you estimate progress.
Watch for file size limits. If individual files exceed 25 MB, split them before uploading. Large CSV files can be split by row count using a spreadsheet application. Large log files can be split using a text editor.
Process different systems separately. If you are consolidating from multiple old systems, extract from each system's files as a separate project. This makes it easier to trace where addresses came from and to identify which system's data is causing issues if problems arise during import.