Batch Email Extraction: Tips for Processing Large File Sets
On this page
When You Need Batch Extraction
Single-file extraction is straightforward: upload a file, extract, download. But many real-world tasks involve dozens or hundreds of files. Consolidating contacts from a company archive, migrating between platforms, or processing records from multiple departments all require working through files in batches.
Email Extractor accepts up to 25 MB per file and 100 MB per file batch. This guide covers how to plan and execute extraction when you have more files than a single batch can handle.
Planning Your Batches
Assess your files first
Before starting, survey what you have:
- Count and total size. How many files, and how large are they combined? This determines how many batches you will need.
- File types. Which formats are in the set? Email Extractor supports 19 file types: TXT, TSV, CSV, HTML, HTM, XML, JSON, MD, LOG, VCF, DOCX, XLSX, XLSM, XLSB, XLS, ODS, PDF, MSG and EML. Any unsupported formats need preprocessing. MBOX files need to be converted to EML. ZIP archives need to be unzipped.
- Files over 25 MB. Any single file over 25 MB needs to be reduced before uploading. Large CSVs can be split by row count. Large log files can be split with a text editor. Large spreadsheets can be split across multiple workbooks.
Choose a batching strategy
There are several ways to group files into batches:
By source or system. Group files that came from the same system together. All exports from your CRM in one batch, all saved emails in another, all PDF contracts in a third. This makes it easier to trace where extracted addresses came from.
By file type. Group all XLSX files in one batch, all PDFs in another, all EML files in a third. This can help if you want to handle format-specific issues (like scanned PDFs needing OCR) separately.
By size. Simply fill each batch up to the 100 MB limit. This is the fastest approach when provenance tracking is not important.
By date or project. Group files by time period or project. This is useful when you want to keep extracted lists aligned with specific campaigns, quarters or business units.
Running the Extraction
Step-by-step for each batch
- Go to Email Extractor.
- Select "Text and files."
- Upload the files for this batch (up to 100 MB total).
- Click "Extract emails."
- Review the results.
- Download as CSV with sources.
The CSV with sources format is especially useful for batch work. The source column records which file each address was found in. When you are processing many files across many batches, this provenance trail helps you trace any address back to its origin.
Keep a processing log
For large file sets, track your progress:
- Which files have been processed.
- How many addresses each batch produced.
- Any files that failed or needed preprocessing.
- The filename of each downloaded results file.
A simple spreadsheet works well for this. Columns for batch number, files included, address count and download filename keep everything organised.
Deduplicating Across Batches
Email Extractor removes duplicate addresses within each extraction run. However, if the same address appears in files from different batches, it will appear in multiple result downloads.
To deduplicate across batches:
- Open all of your downloaded CSV files.
- Copy the address columns into a single text block.
- Paste the combined text into Email Extractor using the text input.
- Click "Extract emails."
The tool applies case-insensitive deduplication, so user@example.com and User@Example.com are treated as the same address. The result is your consolidated, deduplicated master list.
If you need to preserve the source information from individual batches, keep your original per-batch CSV downloads as a reference. The deduplicated run gives you the clean list; the per-batch files give you provenance.
Handling Common Issues
Files that produce no results
If a file produces no extracted addresses, check:
- Is it a scanned PDF? Scanned documents are images, not text. Email Extractor does not include OCR. Process the file through an OCR tool first. See Extract emails from PDF files.
- Is the file actually empty? Some exports from systems produce files with headers but no data rows.
- Are the addresses in an unusual format? Email Extractor uses pattern matching to identify email addresses. If addresses are deliberately obfuscated (like
user [at] example [dot] com), they will not be detected. See Email obfuscation explained. - Is it a supported format? Check that the file extension matches one of the 19 supported types. Renaming an unsupported format's extension does not convert it.
Files over 25 MB
Individual files over the 25 MB limit need to be split or reduced:
- CSV or TSV files: open in a spreadsheet application and save subsets of rows as separate files.
- Log or TXT files: split using a text editor or a command-line tool that divides by line count.
- XLSX files with multiple sheets: save each sheet as a separate file.
- PDF files: use a PDF editor to split into smaller page ranges.
Batches approaching the 100 MB limit
The 100 MB batch limit applies to the combined file size of all files uploaded in a single extraction. If you are close to the limit, leave some headroom. Upload slightly fewer files per batch rather than trying to hit exactly 100 MB, since file sizes reported by your operating system may differ slightly from how the browser calculates them.
Mixed file types in one batch
Email Extractor handles mixed file types in a single upload. You can upload a PDF, three XLSX files and ten EML files together in one batch. There is no need to separate file types unless you want to for organisational reasons.
Optimising for Speed
Prioritise high-value files
If you have a large set and limited time, start with the files most likely to contain useful addresses:
- Structured data files (CSV, XLSX, VCF) typically have the highest density of valid addresses.
- Email files (EML, MSG) contain addresses in headers and may have additional ones in the body.
- Text and log files vary widely but can contain large numbers of addresses if they are application logs or output files.
- PDF and DOCX files tend to have fewer addresses per file since they are usually individual documents.
Use the browser extension for web sources
If some of your sources are web pages rather than files, the Email Extractor browser extension can process up to 25 same-site pages at a time. Use the extension for web sources and the main tool for file-based sources.
Note that webpage and file results do not automatically merge. You will need to combine and deduplicate them separately.
Process during off-hours
For very large extraction tasks (hundreds of files), consider processing during times when you do not need your browser for other work. Extraction happens in your browser, and processing many large files can be resource-intensive.
After Extraction
Once all batches are processed and deduplicated:
- Validate addresses through an email validation service. Email Extractor does not validate whether addresses are deliverable.
- Remove unwanted addresses like role-based addresses (info@, support@), disposable addresses, and system addresses (noreply@, mailer-daemon@).
- Check against your suppression list. See What is an email suppression list.
- Format for your target platform. See How to format your email list for a CRM.