guide

Email A/B Testing - A 2026 Guide to Subject Lines, Send Times and Real Significance

How to A/B test email properly - subject lines, send times, and content - with the split mechanics, sample-size math, and why you should measure conversions instead of just open rate.

Published:

Email is one of the easiest places to A/B test and one of the easiest to test badly. The mechanics look simple - send version A to half, version B to the other half, pick the higher open rate - and that simplicity hides three ways to reach a confident wrong answer. The discipline that makes email tests trustworthy is the same as any experiment - a real metric, enough sample, and honest significance - applied to a channel where the temptation to call it early is strongest.

This guide covers what to test, how the split actually works, how much list you need, and why the metric you pick decides whether your test helps or quietly hurts. The underlying statistics are universal, so what is A/B testing and statistical significance in A/B testing are the foundation this builds on.

What to test, in priority order

Not all email variables move the needle equally. Test in roughly this order of impact.

  1. Subject line. The highest-leverage variable, because it gates every downstream action - nobody clicks an email they did not open. Test length, personalization, curiosity versus clarity, and emoji versus none.
  2. Send time and day. A large, easy-to-run test with real effect, since inbox timing changes visibility. Watch for timezone effects across a global list.
  3. From name and preview text. Small surface, sometimes surprising lift, since they sit right next to the subject line in the inbox.
  4. Body and CTA. Layout, single versus multiple calls to action, button copy and placement. Lower up-front impact but where clicks and conversions are actually won.

Test one variable at a time. If you change the subject line and the send time in the same variant, a win tells you nothing about which change caused it.

How the split actually works - champion versus challenger

Consumer email tools use a pattern worth understanding even if your platform automates it. Instead of a permanent even split, you run a champion-challenger send:

  1. Hold out a test fraction of the list - say 20 percent.
  2. Split that holdout evenly between variant A and variant B.
  3. Wait a decision window - a few hours for opens, longer if you can wait for clicks.
  4. Send the winning variant to the remaining 80 percent.

This banks most of the upside of the better variant while exposing only a slice of the list to the worse one. The catch is the decision window. A short window picks a winner on early opens before later opens and conversions arrive, which is the email version of the peeking problem - calling a result before the data is in. Our post on sequential testing and the peeking problem explains why early looks inflate false positives, and the same logic applies to a two-hour open-rate readout.

How much list do you need

The honest answer depends entirely on the metric, because sample size scales with how rare the event is.

MetricTypical baselineSample needed per variantNotes
Open rate20 to 40%A few thousandHigh baseline, reaches significance fast
Click rate1 to 5%Tens of thousandsLower baseline, needs far more
ConversionUnder 2%Very large / batch sendsOften only detectable across many campaigns

The pattern is clear - the closer your metric is to the money, the more list you need to detect a change in it. If your list is small, test subject lines on open rate where a few thousand recipients suffice, and be realistic that only large effects will ever clear significance on clicks. To turn a target lift into a required recipient count, how to calculate sample size for A/B testing walks through the exact inputs - baseline rate, minimum detectable effect, and power.

The metric trap - opens are not the goal

The single most common email-testing mistake is optimizing open rate as if it were the objective. It is not. A subject line that boosts opens by baiting curiosity can lower clicks and conversions, because the people it lured in were never going to act. Open rate is a means, conversions are the end, and in 2026 open rate is also noisier than it used to be - privacy features that pre-fetch images fire the tracking pixel whether or not a human read the message, inflating opens for affected recipients.

The practical rule is to test the subject line on opens if that is the lever, but always verify the winning variant did not lose downstream. Attributing real conversions to an email send is where a product-analytics or warehouse tool earns its place. PostHog can tie a send to the on-site events and funnels that follow it, and its experiments read straight from those product metrics. If your conversion and revenue data live in a warehouse, GrowthBook joins experiment assignment to that data directly so the metric you judge on is the money, not the open. Neither sends the email - your ESP does that - but both answer the question that actually matters, which is whether the click became a customer.

Common pitfalls

  • Declaring a winner on opens in two hours. Give the decision window enough time for clicks to land, or you are peeking.
  • Testing multiple variables at once. One change per variant, or the result is uninterpretable.
  • Ignoring list segmentation. A subject line that wins for engaged subscribers can lose for a cold segment. Test within comparable audiences.
  • Treating open rate as truth. Pixel pre-fetching biases it upward. Use it for relative comparison, not as a hard number.

The short version

  • Prioritize subject line and send time - the highest-leverage, easiest-to-run email tests.
  • Use a champion-challenger split but give the decision window enough time for downstream metrics to arrive.
  • Size the test by metric - opens need thousands, clicks and conversions need far more.
  • Do not optimize opens in isolation - verify the winner also holds up on clicks and conversions.
  • Attribute to real outcomes with a product-analytics or warehouse tool so you are judging revenue, not curiosity.

Email rewards the same rigor as any experiment - the channel is just faster and more tempting to call early. Build the habit with our full how to run an A/B test walkthrough, and your subject-line tests will start compounding instead of chasing noise.

Frequently Asked Questions

How big does an email list need to be for A/B testing?

Big enough that a realistic difference is detectable, which depends on the metric. Open rate is a high-baseline metric - often 20 to 40 percent - so subject-line tests reach significance on a few thousand recipients per variant. Click and conversion rates are far lower, often low single digits, so detecting a change in those needs tens of thousands per variant. If your list is small, test the high-baseline metric first, batch several sends together, or accept that only large effects will ever show up as significant.

Should I test on my whole list or a sample first?

The classic email pattern is a champion-challenger send. You hold out a fraction of the list, split that holdout evenly between two variants, wait a set window, then send the winner to the remaining majority. This captures most of the upside of the better variant while limiting the downside of the worse one. The trade-off is that a short decision window can pick a winner on early opens before later opens and, more importantly, conversions have landed.

What should an email A/B test actually measure?

Whatever is closest to the money. Open rate tells you the subject line worked, but a subject line that boosts opens by baiting curiosity can lower clicks and conversions. Click rate is better, and conversions or revenue attributed to the send is best. The honest practice is to test the subject line on opens if that is the lever, but always check that the winning variant did not lose downstream on the metric you actually care about.

Is email open rate a reliable metric in 2026?

Less than it used to be. Privacy features that pre-fetch images inflate opens by firing the tracking pixel whether or not a human read the message, so open rate is noisier and biased upward for affected recipients. That does not make it useless for a relative comparison between two subject lines sent to similar audiences, but it does mean you should not treat open rate as ground truth and should lean on clicks and conversions wherever the sample allows.

Explore More

Free Newsletter

Get the Feature Flags Newsletter

Platform benchmarks, real pricing data and progressive delivery practice. No spam.

Free. Unsubscribe any time. See our privacy policy.

Related Articles