Article content and detailed guides remain in English. The selected language applies to controls and quick instructions.

Back to articles

Extracting Emails from XML Data Feeds, API Responses, RSS Feeds, Sitemaps, SOAP Payloads and Configuration Files

On this page

XML as an Email Data Source

XML (Extensible Markup Language) is a structured data format used across enterprise systems, APIs, data feeds, content syndication and configuration management. Email addresses appear in XML data more often than most organisations realise, embedded in contact records, user profiles, notification configurations, directory listings, content metadata and system configurations:

XML source type Where email addresses appear Typical use case
API responses (REST returning XML, SOAP) User profile elements; contact information elements; notification recipient fields; order data (customer email, billing email, shipping contact) Extracting customer or contact emails from CRM, ERP or e-commerce API exports
RSS and Atom feeds Author elements (managingEditor, webMaster in RSS; author/email in Atom); contributor elements; contact fields in podcast feeds Extracting author and contributor emails from blog, news and podcast feeds
Data exports (CRM, ERP, HR systems) Contact records; employee directories; vendor records; customer records Consolidating email addresses from legacy system exports when migrating to a new platform
SOAP web service payloads Request and response elements containing contact data; notification endpoint configurations Extracting contact data from enterprise web service integrations
Configuration files (web.config, app.config, XML-based configs) Error notification recipients; admin email settings; SMTP configuration; alert recipient lists Auditing which email addresses receive system notifications and alerts
Directory and listing feeds (RESO/real estate, industry data feeds) Agent/broker contact elements; listing contact information; office directory entries Extracting real estate agent, broker or business contact emails from industry data feeds
Sitemaps (sitemap.xml) Typically no email addresses in sitemaps themselves; but sitemaps list URLs whose pages may contain contact information Using sitemaps to identify pages that contain contact information (download pages, then extract from HTML)

How Email Extractor Handles XML

Email Extractor supports XML as one of its 19 supported file types. When you upload an XML file, Email Extractor scans all text content within the XML structure, including element text content, attribute values and CDATA sections, and extracts any string matching an email address pattern. It does not require knowledge of the XML schema, element names or structure:

XML feature How Email Extractor handles it
Element text content (<email>user@example.com</email>) Extracted: the text content of every element is scanned
Attribute values (<contact email="user@example.com">) Extracted: attribute values are scanned
CDATA sections (<![CDATA[Contact: user@example.com]]>) Extracted: CDATA content is treated as text and scanned
Namespaced elements (<ns:email>user@example.com</ns:email>) Extracted: namespaces do not affect text content extraction
Nested elements (<contacts><contact><email>user@example.com</email></contact></contacts>) Extracted: all levels of nesting are scanned
Comments (<!-- admin@example.com -->) Extracted: comment text is scanned
Processing instructions (<?contact email="user@example.com"?>) May or may not be extracted depending on parser behaviour; do not rely on this
Entity references (&amp; etc.) Decoded before scanning; user&#64;example.com is recognised as user@example.com
Binary/encoded content (Base64-encoded attachments within XML) Not decoded: Base64-encoded content is not decoded, so email addresses within encoded attachments are not extracted

What Email Extractor does NOT do with XML

Limitation Explanation Workaround
No XPath or element-specific extraction Email Extractor extracts all email addresses from the entire XML file; it cannot extract only from specific elements (such as only <billing_email> but not <shipping_email>) Extract all, then filter in the resulting CSV using the source information
No schema validation Email Extractor does not validate the XML against a schema; it processes the file as text with XML-aware parsing Validate XML separately before uploading if schema compliance matters
No XSLT transformation Email Extractor does not apply XSLT transformations Transform XML using an XSLT processor first, then upload the result
No recursive URL fetching If the XML file contains URLs (such as a sitemap listing page URLs), Email Extractor does not fetch those URLs Download the pages listed in the sitemap separately, then upload the HTML files
No JSON-in-XML parsing If an XML element contains a JSON string as its text content, Email Extractor treats it as text (which may still match email patterns) For complex nested formats, extract the JSON content and upload it as a separate .json file

Common XML Sources and Extraction Workflows

CRM and ERP data exports

System Export format Email fields Extraction workflow
Salesforce CSV or XML data export (Setup > Data Export) Contact email; lead email; account email; user email; custom email fields Export as XML; upload to Email Extractor; download deduplicated list as CSV with sources to identify which records share email addresses
HubSpot CSV export from contacts, companies, deals Contact email; company email; associated contact emails Export as CSV; upload to Email Extractor; deduplicate
SAP iDoc (XML); CSV exports from SE16/SE38 Business partner email; vendor email; customer email; employee email Export iDocs or transaction data as XML; upload to Email Extractor
Oracle / NetSuite SuiteAnalytics XML; saved search CSV/XML Customer email; vendor email; employee email; transaction contact email Export saved searches as XML; upload to Email Extractor
Legacy systems (custom, older platforms) XML export (if available); CSV; fixed-width text Varies by system Export in whatever format available; if XML, upload directly; if CSV, upload directly; if fixed-width, convert to CSV or TXT first

RSS and Atom feeds

Feed element Standard Example Email Extractor result
RSS managingEditor RSS 2.0 <managingEditor>editor@example.com (Jane Smith)</managingEditor> Extracts editor@example.com
RSS webMaster RSS 2.0 <webMaster>webmaster@example.com</webMaster> Extracts webmaster@example.com
Atom author email Atom 1.0 <author><name>Jane Smith</name><email>jane@example.com</email></author> Extracts jane@example.com
iTunes owner (podcasts) iTunes RSS <itunes:owner><itunes:email>podcast@example.com</itunes:email></itunes:owner> Extracts podcast@example.com
Content body RSS/Atom Email addresses mentioned in article body (encoded in <content:encoded> or <content>) Extracts any email patterns in body text

Configuration file auditing

Config file type Where emails appear Why you would extract them
web.config (ASP.NET) SMTP settings; error notification recipients; custom app settings with email values Audit which addresses receive error notifications; verify configuration after deployment
app.config / appsettings.xml Notification recipients; admin email settings; service account emails System administration; identify all email addresses in application configuration
log4j.xml / logback.xml SMTPAppender recipient and from addresses Audit logging notification recipients; update after staff changes
Jenkins/CI config.xml Admin email; notification recipients; build failure recipients Audit CI/CD notification configuration; update after team changes
WordPress wp-config and plugin XML Admin email; plugin notification settings Site administration; identify all configured notification addresses

Batch processing multiple XML files

Scenario Workflow
Multiple API response files Save each API response as a separate .xml file; upload all files at once to Email Extractor (batch upload); download results as CSV with sources to see which email came from which API response file
Daily data feed processing Save each day's data feed as a dated XML file (feed-2026-10-11.xml); upload the batch of files; CSV with sources shows which feed each email appeared in
Multi-system consolidation Export XML from each system (CRM, ERP, HR, support); upload all XML files in one batch; deduplicated results show unique emails across all systems; source column shows which systems each email appears in
Configuration audit across servers Collect configuration XML files from all servers; upload batch; results show all email addresses configured across the environment; identify addresses that need updating (departed employees, old distribution lists)

File Size and Format Considerations

Consideration Detail
Maximum file size Email Extractor supports files up to 25 MB each and 100 MB per batch
Large XML files If your XML export exceeds 25 MB, split it into smaller files before uploading; most XML exports can be split by record (each record as a separate file) or by section
Encoding Email Extractor handles UTF-8 and common encodings; if your XML uses an unusual encoding, verify that email characters (@ and domain names) are correctly represented
Compressed XML Email Extractor does not process compressed files (ZIP, GZIP); decompress XML files before uploading
XML with embedded binaries If your XML contains Base64-encoded attachments (common in SOAP and email XML formats like EML), those encoded sections are not decoded; email addresses within encoded attachments are not extracted
Malformed XML Email Extractor processes the file as text if XML parsing fails, so email addresses are still extracted even from malformed XML, though element boundaries may not be accurately determined

Extract emails

Explore tools

Verify emails

Check address validity before using your list.

ZeroBounce

Email Verification

Verifies email lists and provides tools for monitoring deliverability.

Useful when list cleaning and sender health belong in one workflow.

Explore ZeroBounce (opens in a new tab)