A/B Testing Cold Emails: What to Test and How to Measure Results
On this page
Why A/B Test Cold Emails
Most cold email campaigns start with guessing: you write a subject line that sounds good, a body that seems persuasive, and hit send. Sometimes it works. Often it does not, and you do not know why.
A/B testing replaces guessing with data. By sending two versions of the same email to comparable groups and measuring the results, you learn what actually drives opens, replies and meetings. Over time, each test compounds into a better-performing campaign.
How Cold Email A/B Testing Works
The basic structure
- Create two versions of your email that differ by exactly one element (the variable you are testing).
- Split your prospect list into two equal, random groups.
- Send Version A to Group A and Version B to Group B.
- Wait for results to accumulate (typically 3-5 days for cold email).
- Compare the performance metric that matches your variable.
- Use the winner as your new baseline. Test the next variable.
One variable at a time
This is the most important rule. If you change the subject line and the opening line and the CTA at the same time, you cannot know which change caused the difference in results.
Test one element per experiment:
- Subject line only.
- Opening line only.
- Body copy only.
- Call to action only.
- Send time only.
- Sender name only.
What to Test
1. Subject lines
Subject lines determine whether your email gets opened. This is the highest-impact element to test.
Variables to test:
- Length. Short (3-5 words) vs medium (6-8 words) vs long (9+ words).
- Personalisation. Including the prospect's name or company vs not including it.
- Question vs statement. "Struggling with X?" vs "A better approach to X."
- Specificity. "Improve your email deliverability" vs "How Company X reduced bounces by 40%."
- Curiosity. Vague ("Quick question") vs direct ("Re: your hiring post").
- Case. Title Case vs Sentence case vs all lowercase.
Metric: Open rate.
For subject line ideas, see Cold Email Subject Lines That Get Opened.
2. Opening lines
The opening line appears in the email preview alongside the subject line. It determines whether the prospect reads beyond the first sentence.
Variables to test:
- Personalisation depth. Generic industry reference vs specific company mention vs reference to the prospect's recent activity.
- Problem-first vs compliment-first. Leading with a pain point vs leading with something positive about the prospect or their company.
- Direct vs contextual. Getting straight to the point vs establishing why you are reaching out.
Metric: Reply rate (opens stay constant if the subject line is identical).
3. Body copy
The email body builds the case for your ask.
Variables to test:
- Length. Short (50-75 words) vs medium (100-150 words) vs long (200+ words).
- Social proof. Including a case study reference vs not including one.
- Value proposition framing. Feature-focused vs outcome-focused.
- Number of points. One clear benefit vs three bullet points.
- Tone. Formal vs conversational.
Metric: Reply rate or click rate (depending on whether your CTA is a reply or a link).
4. Call to action (CTA)
The CTA is what you ask the prospect to do.
Variables to test:
- Ask type. "Would you be open to a 15-minute call?" vs "Want me to send over a case study?" vs "Does this sound relevant to your team?"
- Specificity. "Can we chat sometime?" vs "Are you free Thursday at 2pm?"
- Commitment level. Low commitment ("I can send more info") vs higher commitment ("Let's book a call").
- Question vs link. Asking for a reply vs including a calendar link.
Metric: Reply rate or click rate.
5. Send time
When you send affects whether your email is seen.
Variables to test:
- Day of week. Tuesday vs Thursday. Weekday vs weekend.
- Time of day. Early morning (7-8am) vs mid-morning (10-11am) vs afternoon (2-3pm) vs evening (6-7pm).
- Time zone targeting. Sending at 9am in the prospect's local time zone vs a fixed time.
Metric: Open rate and reply rate.
6. Sender name and email
Who the email appears to come from affects trust and open rates.
Variables to test:
- Full name vs first name. "Sarah Johnson" vs "Sarah."
- Name plus company. "Sarah from Acme" vs "Sarah Johnson."
- Title inclusion. Including your job title in the sender name vs not.
- Sender role. Sending from a founder vs a sales rep vs a customer success manager.
Metric: Open rate.
Setting Up the Test
Sample size
You need enough recipients in each group to get reliable results. With cold email, response rates are typically low (1-10%), so you need larger samples than you might expect.
Rule of thumb: At least 100 recipients per variant for open rate tests. At least 200-300 per variant for reply rate tests (because reply rates are lower than open rates, you need more data to detect a meaningful difference).
If your list is too small, you will not be able to distinguish a real difference from random variation.
Randomisation
Split your list randomly, not alphabetically or by any other attribute. If Group A gets all the tech companies and Group B gets all the healthcare companies, your results measure industry differences, not email differences.
Most cold email platforms handle randomisation automatically when you set up an A/B test.
Control for variables
Both groups should be as similar as possible in every way except the one variable you are testing:
- Same sending domain and mailbox.
- Same day and time (unless testing send time).
- Same follow-up sequence (unless testing the follow-up).
- Similar company sizes, industries and geographies.
Run duration
Wait long enough for results to stabilise. Cold email responses can trickle in over several days.
Recommended: Run each test for at least 5-7 business days before declaring a winner. Some prospects reply to the first email days after receiving it.
Measuring Results
Key metrics
| Testing | Primary metric | Secondary metric |
|---|---|---|
| Subject line | Open rate | Reply rate |
| Opening line | Reply rate | Positive reply rate |
| Body copy | Reply rate | Meeting booked rate |
| CTA | Reply rate / Click rate | Meeting booked rate |
| Send time | Open rate | Reply rate |
| Sender | Open rate | Reply rate |
Statistical significance
A result is statistically significant when it is unlikely to have occurred by chance. For cold email testing, aim for at least 90% confidence (95% is standard in academic research, but 90% is practical for cold email optimisation).
Example: Version A gets 22% open rate and Version B gets 28% open rate. With 100 emails per group, the difference might be random noise. With 500 per group, the same difference is much more likely to be real.
Use an online A/B test significance calculator. Enter the sample sizes and conversion rates for each variant. If the result is significant, adopt the winner. If not, either run the test longer or accept that the difference is too small to matter.
What counts as meaningful
Not every statistically significant difference is worth acting on.
- Open rate difference of 5+ percentage points: Worth investigating.
- Reply rate difference of 1-2+ percentage points: Meaningful for cold email, where baseline reply rates are low.
- A consistent winner across multiple tests: More reliable than a single test.
Building a Testing Programme
Start with subject lines
Subject lines have the highest impact and are the easiest to test. Run 2-3 subject line tests before moving to other variables.
Keep a log
Document every test:
- Date.
- Variable tested.
- Version A description.
- Version B description.
- Sample size per group.
- Results (with dates and significance).
- Winner and why.
This log prevents re-testing things you have already learned and builds institutional knowledge.
Apply learnings to sequences
Once you find a winning subject line or opening line, apply it across your entire sequence, not just the email you tested. Patterns that work for the first touch often work for follow-ups too.
Re-test periodically
What works changes over time. Subject lines that performed well six months ago may have become overused. Re-test your best performers every few months.
Common Mistakes
Testing too many things at once
Changing three variables and declaring the combined version the winner teaches you nothing. Test one variable at a time.
Declaring winners too early
A 24-hour test with 50 emails per group is not reliable. Wait for sufficient sample size and duration.
Ignoring the downstream metric
A subject line that gets 40% open rate but 0.5% reply rate is worse than one that gets 25% open rate and 3% reply rate. Track the metric that matters most to your business: replies, meetings or revenue.
Not segmenting
What works for enterprise prospects may not work for SMBs. What works for marketing leaders may not work for engineering leaders. Test within segments, not across your entire list.