Extract episode pages (title, date, guest name, show notes)
Python (Beautiful Soup, Scrapy); Playwright
4
Parse guest information from show notes
Text parsing; regex; NER (named entity recognition)
5
Extract links from show notes (LinkedIn, company website, personal site)
HTML link extraction
6
Compile guest list with all available fields
Python (pandas); spreadsheet
From RSS feeds (XML)
Step
Action
Tools
1
Find podcast RSS feed URL
Podcast website; Apple Podcasts listing; Podchaser
2
Parse RSS feed
Python (feedparser library)
3
Extract episode data (title, description, publication date, links)
XML parsing
4
Parse guest names from episode titles and descriptions
Text parsing; regex; pattern matching
5
Extract URLs from description HTML
HTML link extraction
6
Compile guest list
Python (pandas); spreadsheet
From podcast directories (API)
Step
Action
Tools
1
Search for podcasts by keyword/topic using API
Listen Notes API; Podchaser API
2
Retrieve episode listings for each podcast
API requests
3
Extract guest information from episode metadata
JSON parsing
4
For each guest: find email via company website or email finder
Email finding tools; company website lookup
5
Verify and deduplicate guest list
Verification API; deduplication
Common Extraction Patterns
Guest name extraction from episode titles
Pattern
Example
Extraction approach
"Guest Name - Topic"
"Sarah Chen - Building a Remote Team"
Split on " - "; first part is guest name
"Topic with Guest Name"
"Building a Remote Team with Sarah Chen"
Pattern: "with [Name]" at end
"Ep. XX: Guest Name on Topic"
"Ep. 42: Sarah Chen on Remote Work"
Pattern after episode number; before "on"
"#XX Guest Name: Topic"
"#42 Sarah Chen: How to Build Remote Teams"
Pattern after number; before colon
"Topic (Guest: Name)"
"Remote Work (Guest: Sarah Chen)"
Pattern inside parentheses
"Topic feat. Guest Name"
"Remote Work feat. Sarah Chen"
Pattern after "feat." or "ft."
Guest bio extraction from show notes
Field
Common patterns
Notes
Title and company
"Sarah Chen is the CEO of ExampleCorp"
Usually first sentence of bio
LinkedIn
"Connect with Sarah on LinkedIn: [link]"
Look for linkedin.com URLs
Twitter / X
"Follow Sarah: @sarahchen" or Twitter link
Look for twitter.com or x.com URLs
Website
"Learn more at examplecorp.com" or link
Domain URLs in bio
Email
Rare in show notes
More often found via company website
Using Podcast Guest Data for Outreach
Personalisation approach
Element
What to include
Example
Reference the episode
"I listened to your episode on [podcast name]..."
Shows genuine engagement
Reference a specific point
"Your point about [specific insight] resonated because..."
Proves you actually listened
Connect to your offering
"You mentioned [challenge]. That is exactly what [your product] helps with..."
Natural bridge to value proposition
Timing
Reach out within 1-2 weeks of episode publication
Episode is still top of mind
Cold email template (podcast-triggered)
Section
Content
Subject line
"Your [Podcast Name] episode on [topic]"
Opening
"I caught your episode on [Podcast Name] last week. Your perspective on [specific point] was spot on..."
Connection
"You mentioned [challenge/trend/insight]. We have been seeing the same thing with our customers in [their industry]..."
Value
"We help [their role type] solve [that challenge] by [brief value prop]"
CTA
"Would it make sense to compare notes? I think there might be some useful overlap"
Challenges
Challenge
Why it happens
Solution
Guest names not in structured format
Show notes are free-text; inconsistent formatting
Use pattern matching; manual review
No email in show notes
Hosts rarely share guest email addresses
Use email finding tools; company website lookup
Duplicate guests across podcasts
Prolific guests appear on many shows
Deduplicate by name + company
Outdated information
Guest may have changed companies since recording
Cross-reference with LinkedIn; verify before outreach
Episode behind paywall
Some premium podcast content is not publicly accessible
Focus on free / public episodes
Dynamic / JavaScript-loaded content
Modern podcast sites load content dynamically
Use headless browser (Playwright)
Processing Podcast Guest Data
After extracting guest information from podcast show notes (HTML), RSS feeds (XML), directory listings and YouTube descriptions, upload the compiled files to Email Extractor to extract and deduplicate email addresses found in guest bios and show notes. Prolific guests appear across dozens of podcast episodes, so deduplication prevents contacting the same person multiple times from different episode extractions.