Email Extraction for Journalists: Sources, Press Contacts and FOIA Documents
On this page
Why Journalists Extract Email Addresses
Journalists work with large volumes of documents that contain contact information: press releases, court filings, FOIA responses, leaked documents, government directories, corporate filings and event press lists. Finding the right contact buried in a stack of PDFs takes time that could go toward reporting.
Email Extractor processes these documents in bulk and returns every email address found. Because the tool runs in your browser, sensitive source material stays on your computer.
Common Sources
Press releases and media advisories
PR agencies and corporate communications teams distribute press releases as PDFs, Word documents or plain text. These typically include a media contact section with email addresses and phone numbers. When you accumulate dozens of press releases for a beat, extracting from all of them builds a press contact database over time.
FOIA and public records responses
Freedom of Information Act requests and public records responses often arrive as PDFs, sometimes hundreds of pages. Government emails, internal correspondence and agency communications within these documents contain addresses of officials, staff and external contacts involved in the matter.
Upload FOIA PDFs to Email Extractor to identify every person whose address appears in the documents. Download as CSV with sources to trace each address to its source page.
Text-based PDFs work directly. Scanned FOIA documents (common from agencies that print and re-scan) require OCR before extraction.
Court filings and legal documents
Court documents, complaints, motions and settlement agreements contain attorney contact details, party information and sometimes witness or expert contacts. These are typically available as PDFs from PACER or state court filing systems.
Government staff directories
Federal, state and local government agencies publish staff directories on their websites or make them available as downloadable files. Save these as HTML or download the files and extract addresses for contacting specific departments or officials.
Corporate filings and SEC documents
Annual reports (10-K), proxy statements (DEF 14A) and other SEC filings contain email addresses for investor relations, legal counsel and corporate officers. These filings are available as PDFs and HTML from the SEC's EDGAR system.
Event and conference press lists
Press credentials for events, conferences and government briefings come with attendee lists or media contact sheets. Extract from these to build a network of fellow journalists and press officers covering your beat.
Source correspondence
If you export email threads with sources as EML or MSG files, Email Extractor can compile addresses from across those conversations. This is useful for identifying everyone involved in an email chain referenced in a story.
Example
An investigative journalist receives a 200-page FOIA response as a text-based PDF. After uploading:
director@example.com
policy.analyst@example.org
external.contractor@example.net
The CSV with sources output traces each address to the PDF:
email,source
director@example.com,foia-response-2026.pdf
policy.analyst@example.org,foia-response-2026.pdf
external.contractor@example.net,foia-response-2026.pdf
The journalist now has a list of every person whose email address appears in the documents, which helps identify key figures and potential sources.
Practical Uses
Source identification
In document-heavy investigations, extracting addresses from the full document set reveals who was involved in the communications. This is faster than reading every page looking for contact details.
Beat contact management
Journalists covering a specific beat (city hall, technology, healthcare) accumulate contacts across press releases, event lists and correspondence. Periodic extraction consolidates these into a searchable list.
Tip line and submission management
Newsrooms that receive tips and submissions via email can export those messages and extract sender addresses to track submission volume, identify repeat tipsters (with appropriate editorial safeguards), or follow up on unresolved tips.
Cross-referencing
Extract addresses from two different document sets and compare them to find overlap. If the same address appears in both a lobbying disclosure and an agency's internal emails, that connection may be worth investigating.
Source Protection Considerations
Sensitive documents
When working with leaked documents, whistleblower submissions or confidential source materials, be aware that extracting and storing email addresses from these documents creates a record that could identify sources if the file were accessed by others.
Email Extractor does not store or transmit your files. However, the extracted address list is on your computer once you download it. Store sensitive extraction results with the same security measures you apply to the source documents themselves.
Metadata in source files
Documents may contain metadata (author names, revision history, creation dates) that is separate from the visible text. Email Extractor scans for email patterns in the content it reads, which may include some metadata fields depending on the file format. Be aware that extracted addresses might come from metadata as well as visible text.
Newsroom policies
Check whether your newsroom has policies about using automated tools on sensitive documents. Some organisations require that certain categories of documents be handled only on air-gapped computers or within specific security protocols.
Troubleshooting
Redacted FOIA documents
Government agencies sometimes redact email addresses in FOIA responses. Redacted text is removed from the document and cannot be recovered by extraction. If addresses are redacted, you may need to submit a follow-up request or use other public sources.
Scanned PDFs
Many FOIA responses are scanned images rather than text-based PDFs. Email Extractor cannot read text from images. Run OCR on the scanned pages first using a tool like Adobe Acrobat, Tesseract or a similar OCR application. See the PDF guide.
Large document sets
Investigations can involve thousands of pages across many files. Upload in batches (up to 100 MB per batch, 25 MB per file) and merge the results.