Skip to content
XemGet started
Back to the journal

Analytics · 4 minute read

Email A/B testing: ask a better question before picking a winner

Plan useful email A/B tests with a clear hypothesis, fair audience split, meaningful success metric, and enough patience to avoid noisy conclusions.

Xem editorial

Split cream and lavender cover showing two different email variants marked A and B under One question. Two possibilities.
Original artwork by Xem editorial

“Which subject line wins?” sounds like a useful question. Often it is too small.

A test becomes more valuable when it asks why someone might respond: does a practical promise help more than a curiosity-led opening? Does a product screenshot clarify the action better than a decorative image? Does a shorter explanation remove friction, or remove information people need?

The winner matters for one campaign. The explanation can improve the next ten.

Write the hypothesis in plain language

A hypothesis connects a change to a reason. “Version B will get more clicks” predicts an outcome but does not explain anything.

Try: “Readers are unsure what the button opens. A destination-specific label will increase useful clicks compared with ‘Learn more.’”

Now the test has a clear variable, a plausible mechanism, and an outcome. Keep the destination identical so the comparison remains about the label.

For a newsletter, another hypothesis might be that a concrete subject line helps readers recognize the edition's relevance. Compare “Three fixes for your welcome email” with a curiosity-led alternative that the content genuinely supports. Avoid a misleading subject simply because it might attract attention.

Change one meaningful variable

If the subject line, send time, layout, and offer all change, you have two campaigns rather than a clean test of one idea.

Choose one variable when you want a result that is easy to interpret. More complex experiments are possible, but they need more planning, traffic, and analysis than many small email programs can justify.

Keep other important conditions consistent: sender identity, audience eligibility, destination, tracking setup, and observation window. If one variant accidentally reaches a more engaged segment, the result may reflect the audience rather than the creative.

AI is useful for generating contrasting approaches. It should not declare which one will perform best. Ask it to explain the difference between options, then choose the test based on a real uncertainty your team has.

Pick the metric before sending

Use an outcome that matches the hypothesis. If the test concerns a button's clarity, clicks to the intended destination are more relevant than opens. If it concerns whether a guide helps people activate a feature, the feature action may be the important outcome.

Define the denominator. “Unique clicked messages divided by accepted messages” is different from “total click events divided by recipients.” Repeated clicks and security scanners can affect the interpretation.

Apple's Mail Privacy Protection is another reason not to rely on small open-rate differences as proof of better subject lines. The metric is an observation with known limits, not a direct count of people paying attention.

Choose a review window before launch. Do not stop a test simply because the preferred version is ahead after an hour.

Split a comparable audience

Random assignment within one eligible audience is a better starting point than sending one version to your newest subscribers and another to your oldest.

Make the groups mutually exclusive. Preserve consent and suppression rules in both groups. Avoid resending the other variant to people who already received the first unless that is a separate, deliberate experiment.

These are general testing principles. This guide does not claim Xem currently has an automatic randomized A/B winner workflow. Use a platform that supports the experiment you need, or prepare auditable, mutually exclusive groups with appropriate tooling and review them before sending.

A manual process needs clear records: which addresses were eligible, how assignment happened, which version they received, and which observation period you used.

Know when the result is mostly noise

Suppose two fictional groups of 100 accepted messages produce five and seven clicked messages. Seven is larger than five. It is not, by itself, strong evidence of a repeatable improvement.

There is no universal minimum sample size. It depends on the baseline rate, the smallest effect worth detecting, and the confidence and power you require. Use a suitable statistical method or calculator before the send, and document its assumptions.

If your audience is small, focus on large, meaningful differences and combine the numbers with qualitative evidence. Replies and usability feedback can reveal a confusing message even when the experiment is too small for a confident statistical conclusion.

“No clear difference” is a useful result. It can stop the team from spending another week debating a tiny copy change.

Keep an experiment notebook

Record the hypothesis, versions, audience rule, assignment method, send time, primary metric, review window, and result. Include known problems such as a broken link or a tracking outage.

Then write one sentence about what the test changes in your process. If the result is inconclusive, say that. If it applies only to a specific audience or offer, preserve that limit.

Avoid turning one result into a permanent rule such as “short subject lines always win.” Email performance depends on context. A better lesson might be: “For this tutorial series, naming the task made the subject easier to understand.”

Your next test should follow from an unresolved question, not the need to fill an experiment calendar.

Sources and further reading

Written with AI assistance and checked against the linked sources. Examples are illustrative. Our editorial approach

email A/B testingsubject line testingemail campaign optimization

Make your next email a good one.

Start with Xem