How to Extract Email Addresses from JSON, XML and Log Files
On this page
When These Formats Are Useful
JSON, XML and log files frequently contain email addresses in application data, API responses, configuration files, exported records and system logs.
JSON is used by API exports, Slack workspace exports, Trello board exports, Notion data and many web application backups.
XML appears in RSS feeds, sitemap files, data interchange formats, SOAP responses and older application exports.
Log files (.log) contain server access logs, application logs, email server logs and debug output where email addresses appear in request data, user actions or error messages.
Email Extractor reads all three formats and scans them for email address patterns.
Extraction Steps
- Open Email Extractor.
- Select Text and files.
- Upload your
.json,.xmlor.logfile, or drag it into the upload area. Maximum 25 MB per file, 100 MB per batch. - Click Extract emails.
- Review the results. Duplicates are removed using case-insensitive matching.
- Copy individual addresses or download as TXT, CSV or CSV with sources.
You can upload all three file types together in a single batch.
JSON: What Is Scanned
For valid JSON, Email Extractor scans string values, including values nested in objects and arrays. It does not scan property names (keys). If parsing fails, it scans the raw text instead.
Example JSON:
{
"users": [
{
"name": "Yuki Tanaka",
"email": "yuki@example.com",
"notes": "Contact priya@example.org for project handover"
},
{
"name": "admin@example.net",
"email": "admin@example.net"
}
]
}
Extracted addresses:
- yuki@example.com (from the
emailvalue) - priya@example.org (from the
notesvalue) - admin@example.net (from the
emailvalue)
The key name "email" is not scanned. The string "admin@example.net" in the name field is scanned because it is a string value, even though it is stored under a name key.
Encoded characters in JSON
For valid JSON, the reader parses the document before scanning its string values. Standard JSON escapes such as \u0040 for @ are decoded, so "yuki\u0040example.com" can yield yuki@example.com. Malformed JSON falls back to scanning its raw text; escapes in that fallback are not decoded as JSON.
XML: What Is Scanned
Email Extractor scans the text content and attribute values of XML elements.
Example XML:
<contacts>
<person name="Yuki Tanaka">
<email>yuki@example.com</email>
<notes>See also priya@example.org</notes>
</person>
<person name="admin@example.net">
<email>admin@example.net</email>
</person>
</contacts>
Extracted addresses:
- yuki@example.com (element text)
- priya@example.org (element text)
- admin@example.net (attribute value and element text, deduplicated)
HTML entities in XML
The pattern matcher decodes numeric character references, including @ and @ for @, plus a small set of named entities such as @, . and &. This is not a complete XML entity parser: custom entity declarations and other named entities need separate handling.
Log Files: What Is Scanned
Log files are read as plain text. Email Extractor scans each line for email address patterns, regardless of the log format.
Example log content:
2026-03-14 09:12:33 INFO User login: yuki@example.com from 192.168.1.10
2026-03-14 09:13:01 WARN Failed auth for priya@example.org. invalid token
2026-03-14 09:14:22 INFO Password reset requested by yuki@example.com
2026-03-14 09:15:00 ERROR Bounce notification for admin@example.net
Extracted addresses:
The duplicate yuki@example.com from the password reset line is removed.
Log-specific considerations
Log files can be large. Files up to 25 MB are supported. For larger logs, split the file or extract the relevant section before uploading.
Log entries often contain timestamps, IP addresses and other data alongside email addresses. These do not interfere with extraction. IP addresses do not match email patterns and are not included in results.
Limitations
Valid JSON: string values only. Property names are not scanned after successful parsing. Malformed JSON is scanned as raw text, so text in apparent keys may also match.
Structure and output. The JSON reader traverses nested objects and arrays and records the path of each string value as a source location. XML is scanned as text. Neither format exports complete records with their other fields; the email output is a flat list of unique addresses, with source details available separately.
Encoding limits. Valid JSON escapes and supported character references are decoded. Double-encoded strings, custom XML entities and malformed JSON may need separate preprocessing.
No JavaScript or XSLT execution. XML files with XSLT stylesheets or linked scripts are not processed. Only the static text content of the uploaded file is scanned.
HTML files. If your XML file is actually an HTML document, Email Extractor supports HTML and HTM files directly. Uploaded HTML files do not have their JavaScript executed or linked assets fetched.
Troubleshooting
Missing addresses in JSON
- Open the file in a text editor and search for the missing address. If it appears only as a property name (key), it will not be extracted.
- Confirm that the JSON is valid. Standard escapes in string values are decoded after parsing; malformed or double-encoded data may need correction in a copy of the source.
- If the JSON is minified (no line breaks or whitespace), extraction still works. The tool scans the text regardless of formatting.
Missing addresses in XML
- Numeric @ references (
@,@) are decoded. Check for unsupported named entities or custom entity declarations if an address is missing. - If addresses are inside CDATA sections, they should still be found as part of the text scan.
Large files
JSON and XML exports from applications like Slack or Trello can be several megabytes. Files up to 25 MB are supported. For larger exports, check whether the application offers options to export in smaller segments (by date range or channel, for example).
Source review
When uploading multiple JSON, XML or log files in one batch, download results as CSV with sources to trace each address back to its source file.
Related Guides
- Extract emails from a Slack export. for Slack's JSON export format.
- Extract emails from TXT and TSV files. for other plain-text formats.
- Extract emails from ZIP archives. when data exports arrive as compressed files.
- Extract emails from multiple file types at once. for mixed batches.