How to Extract Email Addresses from a PDF
On this page
When PDF Extraction Is Useful
PDF files commonly contain email addresses in reports, invoices, conference programmes, attendee lists, government filings, academic papers and company directories. Because PDFs are designed for display rather than data exchange, the addresses are embedded in the document text rather than structured as a contact list.
Email Extractor reads available text and email annotations from PDF files. If the PDF contains selectable text (you can highlight words when viewing it), extraction should work. If the PDF is a scanned image, a separate OCR step is needed first.
Extraction Steps
- Open Email Extractor.
- Select Text and files.
- Upload your PDF file or drag it into the upload area. Maximum file size is 25 MB, with a 100 MB batch limit when uploading multiple files.
- Click Extract emails.
- Review the results. Duplicates are removed using case-insensitive matching.
- Copy individual addresses or download as TXT, CSV or CSV with sources.
Example
A file called conference-programme.pdf contains the following text across its pages:
Keynote Speaker: Dr Yuki Tanaka. yuki.tanaka@example.com
Panel Moderator: Anika Shah (anika@example.org)
For registration queries, contact events@example.net
Speaker enquiries: yuki.tanaka@example.com
After extraction, Email Extractor returns three unique addresses:
The duplicate yuki.tanaka@example.com from the speaker enquiries line is removed automatically.
What Email Extractor Reads in a PDF
Text content. The main body text of the PDF, including text in headers, footers, tables and text boxes. This covers any email address that appears as selectable text in the document.
Email annotations. PDF files can contain annotations such as mailto: links. Email Extractor reads these annotations alongside the body text.
Scanned PDFs and OCR
Email Extractor does not include built-in OCR (optical character recognition). If your PDF is a scanned image, such as a photographed document or a scan of a printed page, the file contains image data rather than text. Uploading it will produce no results or incomplete results.
To extract from a scanned PDF:
- Run the PDF through an OCR tool first to convert the images to selectable text. Options include:
- Adobe Acrobat (Edit PDF or Scan & OCR)
- Free online OCR tools (note that these upload your file to a server)
- Open-source tools like Tesseract OCR
- Save the OCR output as a new text-based PDF or as a text file.
- Upload the OCR output to Email Extractor.
How to tell if a PDF is scanned
Open the PDF in any viewer and try to select text with your cursor. If you can highlight individual words, the PDF contains text and extraction should work. If clicking and dragging selects the entire page as an image, the PDF is scanned and needs OCR.
Limitations
Scanned documents. As described above, image-only PDFs require a separate OCR step. Email Extractor cannot read text from images.
Embedded images. If an email address appears only inside an embedded image (such as a business card photo placed in a PDF), it will not be extracted. The address must exist as text in the PDF.
Heavily formatted or layered PDFs. Some design-heavy PDFs use complex text positioning. The visible email address may be split across multiple text elements internally, which can cause partial matches or missed addresses.
Password-protected PDFs. If the PDF requires a password to open, remove the password protection first before uploading.
Troubleshooting
Fewer addresses than expected
- Check whether parts of the PDF are scanned images. A single PDF can contain both text pages and scanned pages. Addresses on scanned pages will not be found.
- Open the PDF and try selecting the text around a missing address. If it cannot be selected, that section is an image.
- Some PDFs use non-standard character encoding for special characters. If an address uses unusual characters, the text extraction may not recognise it.
Addresses appear with extra characters
- PDFs sometimes include invisible characters or formatting marks in the text layer. If an extracted address contains unexpected characters, compare it against the visible text in the PDF.
- Hyperlinked addresses may appear twice in the text layer, once as visible text and once as the link target. Email Extractor removes exact duplicates, but if the two versions differ slightly (extra characters in one), both may appear.
Source review
When extracting from multiple PDFs in one batch, download results as CSV with sources to see which PDF each address came from. This is useful for tracing addresses back to specific documents.
Related Guides
- Extract emails from Word documents. for DOCX files.
- Extract emails from multiple file types at once. for mixed batches of PDFs and other formats.
- Extract emails from business cards using OCR. for scanned card images.