Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

How to Extract Email Addresses from MBOX and Gmail Takeout Exports

On this page

MBOX Files Are Not Directly Supported

MBOX is a standard format for storing collections of email messages in a single file. Gmail Takeout, Thunderbird and other email clients export mailboxes in this format. However, Email Extractor does not read MBOX files directly.

Renaming an MBOX file's extension does not convert it. Changing .mbox to .txt or .eml will not make it work. The file contains multiple concatenated messages and must be split into individual EML files first.

The conversion is straightforward using a short Python script or an email client that can import MBOX and export individual messages.

Step 1: Obtain Your MBOX File

From Gmail via Google Takeout

  1. Go to takeout.google.com.
  2. Click "Deselect all", then scroll down and select Mail.
  3. Click "All Mail data included" to choose specific labels if you want a subset.
  4. Click "Next step", choose your delivery method and file settings.
  5. Create the export and wait for the download link (this can take hours or days for large mailboxes).
  6. Download and unzip the archive. The MBOX file is inside the Takeout/Mail/ folder.

For full instructions, see Google's Takeout documentation.

From Thunderbird

Thunderbird stores mail in MBOX format by default. The files are in your Thunderbird profile folder. Each folder in Thunderbird corresponds to an MBOX file (without a file extension). You can also use the ImportExportTools NG add-on to export specific folders as MBOX files with the .mbox extension.

From other email clients

Many email clients support MBOX export. Check your client's export or backup options for an MBOX or "Unix mailbox" option.

Step 2: Convert MBOX to EML Files

Using Python

This script reads an MBOX file and saves each message as a separate .eml file:

import mailbox
import os

mbox_path = "All mail Including Spam and Trash.mbox"
output_dir = "eml_output"

os.makedirs(output_dir, exist_ok=True)

mbox = mailbox.mbox(mbox_path)
for i, message in enumerate(mbox):
    eml_path = os.path.join(output_dir, f"message_{i:06d}.eml")
    with open(eml_path, "wb") as f:
        f.write(message.as_bytes())

print(f"Saved {i + 1} messages to {output_dir}/")

Save this as convert_mbox.py in the same folder as your MBOX file. Run it with:

python convert_mbox.py

This creates a folder called eml_output containing individual .eml files, one per message.

Python is available by default on macOS and most Linux distributions. On Windows, download it from python.org.

Using Thunderbird

  1. Install Thunderbird if you do not have it.
  2. Install the ImportExportTools NG add-on.
  3. Create a local folder in Thunderbird.
  4. Right-click the folder → ImportExportTools NG → Import MBOX file.
  5. Once imported, right-click the folder → ImportExportTools NG → Export all messages in folder → EML format.

Step 3: Extract from the EML Files

  1. Open Email Extractor.
  2. Select Text and files.
  3. Select multiple .eml files from the output folder. Upload in batches within the 100 MB limit (25 MB per individual file).
  4. Click Extract emails.
  5. Review the results. Duplicates are removed using case-insensitive matching.
  6. Copy individual addresses or download as TXT, CSV or CSV with sources.

For large MBOX files that produce thousands of EML files, work in batches. Upload a group, download the results, then upload the next group. After all batches are processed, merge the result files and remove duplicates.

Example

A Gmail Takeout MBOX file is converted to 500 EML files. The first batch of 50 files is uploaded. Email Extractor scans the headers (From, To, CC) and body text of each message and returns unique addresses.

Downloading as CSV with sources shows which EML file each address came from:

email,source
yuki@example.com,message_000012.eml
priya@example.org,message_000012.eml
newsletter@example.net,message_000045.eml

After processing all batches, the merged list contains every unique address from the entire mailbox.

What Gets Extracted

Email Extractor reads MSG and EML files for:

  • Message headers: From, To, CC and BCC fields.
  • Body text: Addresses mentioned in the message content, signatures and forwarded messages.
  • Not attachments: File attachments within the messages are not read. See the MSG and EML guide for details on extracting from attachments separately.

Filtering Before Conversion

A full Gmail Takeout can be large. If you only need addresses from certain conversations, consider these approaches:

Export specific labels. During the Google Takeout setup, you can select individual Gmail labels instead of "All Mail". Export only the labels relevant to your purpose.

Filter during conversion. Modify the Python script to filter messages by date, subject or sender before saving as EML. This reduces the number of files to upload:

import mailbox
import os
import email.utils
from datetime import datetime

mbox_path = "All mail Including Spam and Trash.mbox"
output_dir = "eml_output"
cutoff = datetime(2025, 1, 1)

os.makedirs(output_dir, exist_ok=True)

mbox = mailbox.mbox(mbox_path)
saved = 0
for message in mbox:
    date_str = message.get("Date", "")
    try:
        date_tuple = email.utils.parsedate_to_datetime(date_str)
        if date_tuple >= cutoff:
            eml_path = os.path.join(output_dir, f"message_{saved:06d}.eml")
            with open(eml_path, "wb") as f:
                f.write(message.as_bytes())
            saved += 1
    except (ValueError, TypeError):
        pass

print(f"Saved {saved} messages from {cutoff.year} onwards.")

Troubleshooting

Python script fails with an encoding error

Some MBOX files contain messages with inconsistent character encoding. If the script fails on a specific message, add error handling:

try:
    with open(eml_path, "wb") as f:
        f.write(message.as_bytes())
except Exception as e:
    print(f"Skipped message {i}: {e}")

This skips problematic messages and continues processing the rest.

Very large MBOX file

Gmail Takeout files can be several gigabytes. The Python script processes the file incrementally (one message at a time), so it handles large files without loading the entire file into memory. The conversion step itself should work regardless of MBOX size.

No addresses found in EML files

  • Open one of the EML files in a text editor. You should see email headers (From, To, Date) at the top. If the file looks like binary data or is empty, the conversion may have failed.
  • Check that the original MBOX file is a valid MBOX and not a different format with the wrong extension.

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)