Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

How to Extract Email Addresses from a PDF

On this page

When PDF Extraction Is Useful

PDF files commonly contain email addresses in reports, invoices, conference programmes, attendee lists, government filings, academic papers and company directories. Because PDFs are designed for display rather than data exchange, the addresses are embedded in the document text rather than structured as a contact list.

Email Extractor reads available text and email annotations from PDF files. If the PDF contains selectable text (you can highlight words when viewing it), extraction should work. If the PDF is a scanned image, a separate OCR step is needed first.

Extraction Steps

  1. Open Email Extractor.
  2. Select Text and files.
  3. Upload your PDF file or drag it into the upload area. Maximum file size is 25 MB, with a 100 MB batch limit when uploading multiple files.
  4. Click Extract emails.
  5. Review the results. Duplicates are removed using case-insensitive matching.
  6. Copy individual addresses or download as TXT, CSV or CSV with sources.

Example

A file called conference-programme.pdf contains the following text across its pages:

Keynote Speaker: Dr Yuki Tanaka. yuki.tanaka@example.com
Panel Moderator: Anika Shah (anika@example.org)
For registration queries, contact events@example.net
Speaker enquiries: yuki.tanaka@example.com

After extraction, Email Extractor returns three unique addresses:

The duplicate yuki.tanaka@example.com from the speaker enquiries line is removed automatically.

What Email Extractor Reads in a PDF

Text content. The main body text of the PDF, including text in headers, footers, tables and text boxes. This covers any email address that appears as selectable text in the document.

Email annotations. PDF files can contain annotations such as mailto: links. Email Extractor reads these annotations alongside the body text.

Scanned PDFs and OCR

Email Extractor does not include built-in OCR (optical character recognition). If your PDF is a scanned image, such as a photographed document or a scan of a printed page, the file contains image data rather than text. Uploading it will produce no results or incomplete results.

To extract from a scanned PDF:

  1. Run the PDF through an OCR tool first to convert the images to selectable text. Options include:
    • Adobe Acrobat (Edit PDF or Scan & OCR)
    • Free online OCR tools (note that these upload your file to a server)
    • Open-source tools like Tesseract OCR
  2. Save the OCR output as a new text-based PDF or as a text file.
  3. Upload the OCR output to Email Extractor.

How to tell if a PDF is scanned

Open the PDF in any viewer and try to select text with your cursor. If you can highlight individual words, the PDF contains text and extraction should work. If clicking and dragging selects the entire page as an image, the PDF is scanned and needs OCR.

Limitations

Scanned documents. As described above, image-only PDFs require a separate OCR step. Email Extractor cannot read text from images.

Embedded images. If an email address appears only inside an embedded image (such as a business card photo placed in a PDF), it will not be extracted. The address must exist as text in the PDF.

Heavily formatted or layered PDFs. Some design-heavy PDFs use complex text positioning. The visible email address may be split across multiple text elements internally, which can cause partial matches or missed addresses.

Password-protected PDFs. If the PDF requires a password to open, remove the password protection first before uploading.

Troubleshooting

Fewer addresses than expected

  • Check whether parts of the PDF are scanned images. A single PDF can contain both text pages and scanned pages. Addresses on scanned pages will not be found.
  • Open the PDF and try selecting the text around a missing address. If it cannot be selected, that section is an image.
  • Some PDFs use non-standard character encoding for special characters. If an address uses unusual characters, the text extraction may not recognise it.

Addresses appear with extra characters

  • PDFs sometimes include invisible characters or formatting marks in the text layer. If an extracted address contains unexpected characters, compare it against the visible text in the PDF.
  • Hyperlinked addresses may appear twice in the text layer, once as visible text and once as the link target. Email Extractor removes exact duplicates, but if the two versions differ slightly (extra characters in one), both may appear.

Source review

When extracting from multiple PDFs in one batch, download results as CSV with sources to see which PDF each address came from. This is useful for tracing addresses back to specific documents.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)