Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Email Extraction for Journalists: Sources, Press Contacts and FOIA Documents

On this page

Why Journalists Extract Email Addresses

Journalists work with large volumes of documents that contain contact information: press releases, court filings, FOIA responses, leaked documents, government directories, corporate filings and event press lists. Finding the right contact buried in a stack of PDFs takes time that could go toward reporting.

Email Extractor processes these documents in bulk and returns every email address found. Because the tool runs in your browser, sensitive source material stays on your computer.

Common Sources

Press releases and media advisories

PR agencies and corporate communications teams distribute press releases as PDFs, Word documents or plain text. These typically include a media contact section with email addresses and phone numbers. When you accumulate dozens of press releases for a beat, extracting from all of them builds a press contact database over time.

FOIA and public records responses

Freedom of Information Act requests and public records responses often arrive as PDFs, sometimes hundreds of pages. Government emails, internal correspondence and agency communications within these documents contain addresses of officials, staff and external contacts involved in the matter.

Upload FOIA PDFs to Email Extractor to identify every person whose address appears in the documents. Download as CSV with sources to trace each address to its source page.

Text-based PDFs work directly. Scanned FOIA documents (common from agencies that print and re-scan) require OCR before extraction.

Court filings and legal documents

Court documents, complaints, motions and settlement agreements contain attorney contact details, party information and sometimes witness or expert contacts. These are typically available as PDFs from PACER or state court filing systems.

Government staff directories

Federal, state and local government agencies publish staff directories on their websites or make them available as downloadable files. Save these as HTML or download the files and extract addresses for contacting specific departments or officials.

Corporate filings and SEC documents

Annual reports (10-K), proxy statements (DEF 14A) and other SEC filings contain email addresses for investor relations, legal counsel and corporate officers. These filings are available as PDFs and HTML from the SEC's EDGAR system.

Event and conference press lists

Press credentials for events, conferences and government briefings come with attendee lists or media contact sheets. Extract from these to build a network of fellow journalists and press officers covering your beat.

Source correspondence

If you export email threads with sources as EML or MSG files, Email Extractor can compile addresses from across those conversations. This is useful for identifying everyone involved in an email chain referenced in a story.

Example

An investigative journalist receives a 200-page FOIA response as a text-based PDF. After uploading:

director@example.com
policy.analyst@example.org
external.contractor@example.net

The CSV with sources output traces each address to the PDF:

email,source
director@example.com,foia-response-2026.pdf
policy.analyst@example.org,foia-response-2026.pdf
external.contractor@example.net,foia-response-2026.pdf

The journalist now has a list of every person whose email address appears in the documents, which helps identify key figures and potential sources.

Practical Uses

Source identification

In document-heavy investigations, extracting addresses from the full document set reveals who was involved in the communications. This is faster than reading every page looking for contact details.

Beat contact management

Journalists covering a specific beat (city hall, technology, healthcare) accumulate contacts across press releases, event lists and correspondence. Periodic extraction consolidates these into a searchable list.

Tip line and submission management

Newsrooms that receive tips and submissions via email can export those messages and extract sender addresses to track submission volume, identify repeat tipsters (with appropriate editorial safeguards), or follow up on unresolved tips.

Cross-referencing

Extract addresses from two different document sets and compare them to find overlap. If the same address appears in both a lobbying disclosure and an agency's internal emails, that connection may be worth investigating.

Source Protection Considerations

Sensitive documents

When working with leaked documents, whistleblower submissions or confidential source materials, be aware that extracting and storing email addresses from these documents creates a record that could identify sources if the file were accessed by others.

Email Extractor does not store or transmit your files. However, the extracted address list is on your computer once you download it. Store sensitive extraction results with the same security measures you apply to the source documents themselves.

Metadata in source files

Documents may contain metadata (author names, revision history, creation dates) that is separate from the visible text. Email Extractor scans for email patterns in the content it reads, which may include some metadata fields depending on the file format. Be aware that extracted addresses might come from metadata as well as visible text.

Newsroom policies

Check whether your newsroom has policies about using automated tools on sensitive documents. Some organisations require that certain categories of documents be handled only on air-gapped computers or within specific security protocols.

Troubleshooting

Redacted FOIA documents

Government agencies sometimes redact email addresses in FOIA responses. Redacted text is removed from the document and cannot be recovered by extraction. If addresses are redacted, you may need to submit a follow-up request or use other public sources.

Scanned PDFs

Many FOIA responses are scanned images rather than text-based PDFs. Email Extractor cannot read text from images. Run OCR on the scanned pages first using a tool like Adobe Acrobat, Tesseract or a similar OCR application. See the PDF guide.

Large document sets

Investigations can involve thousands of pages across many files. Upload in batches (up to 100 MB per batch, 25 MB per file) and merge the results.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)