Skip to main content
CalculatorRuns in your browserFree · No signup

Cold Email A/B Test Calculator

Stop calling a winner on 50 emails. Enter sends and replies for two variants and get a real p-value, confidence intervals, a plain-English verdict and the number of emails you need to send before you can trust the result. Free, no signup.

Test results

Enter emails sent and outcomes for each variant. Everything updates live.

Variant A (control)

Variant B (challenger)

Plan the next test

How many emails per variant you need before starting, at 80% power.

3%

Follows variant A's observed rate until you move it. Use your historical rate when planning from scratch.

20%

A 20% relative lift on 3% means detecting 3.60% or better.

Share this test

All inputs are encoded in the link, so a teammate opens the same result.

Significance

Enter both variants to see the result

Add emails sent and replies for A and B. You'll get rates, the p-value, a confidence interval and a plain-English verdict.

Sample size needed

Per variant

13,914

20% lift on 3%

Total emails

27,828

95% confidence, 80% power

Sample size by minimum detectable effect
Min. detectable liftPer variantTotal
10%53,211106,422
20%13,91427,828
30%6,45512,910
50%2,5185,036
Nothing leaves your browserNo account or email requiredUnlimited use
How it works

How to use the A/B Test Calculator

  1. 01

    Enter both variants

    Pick the metric you are testing (replies, positive replies or opens), then add emails sent and outcomes for variant A and variant B. Choose 90, 95 or 99% confidence.

  2. 02

    Read the verdict

    You get each variant's rate with an interval, the absolute and relative difference, the z statistic, the two-sided p-value and whether the test is significant at your chosen level.

  3. 03

    Plan the sample size

    Set your baseline rate and the smallest lift worth detecting. The planner shows emails per variant at 80% power, plus how many more you need if the current lead is real.

Why 50 emails per variant tells you nothing

Cold email reply rates are small, commonly 1% to 5%, and that is what makes testing hard. If variant A gets 2 replies from 50 emails and variant B gets 4, B looks twice as good. In reality that is a difference of two replies, and the calculator will show a p-value well above 0.05: a gap that size or bigger shows up in a large share of tests where both variants are identical. The lower the base rate, the more emails it takes for a real difference to separate from noise.

Use the sample size planner before you start. At a 3% baseline reply rate, detecting a 20% relative lift (3% to 3.6%) at 95% confidence and 80% power needs about 14,000 emails per variant. Detecting a 50% lift (3% to 4.5%) needs roughly 2,500 per variant, and doubling the rate (3% to 6%) needs about 750. That is why most cold email tests should look for big swings, not polish.

The statistics behind the verdict

The calculator runs a two-proportion z-test. Each variant's rate is outcomes divided by emails sent. Under the null hypothesis that both variants share one true rate, it pools the two samples, computes the standard error of the difference, and expresses the observed difference as a z statistic. The two-sided p-value is the probability of a difference at least that large in either direction if nothing were different. The standard normal CDF is evaluated in the browser with the Abramowitz and Stegun approximation, accurate to about seven decimal places.

  • Significant means p-value below alpha, where alpha is 0.10, 0.05 or 0.01 for 90, 95 or 99% confidence.
  • The confidence interval for the difference uses each variant's own variance. If it excludes zero, the test is significant at that level.
  • Minimum detectable effect (MDE) is the smallest relative lift your sample can reliably detect. Smaller MDE means far more emails, since sample size scales with the inverse square of the difference.
  • 80% power means that if the true lift equals the MDE, four tests in five will reach significance. The remaining fifth will miss a real winner.
  • The tool warns when a variant has fewer than 5 expected outcomes, where the normal approximation gets rough and the p-value should be read as a hint, not a proof.

The peeking problem and what to test

Checking results every day and stopping the moment p dips under 0.05 is the most common way cold email tests go wrong. Random noise crosses that line often on its way to nowhere, so a team that peeks ten times has a much higher false positive rate than the 5% it thinks it has. Fix the sample size with the planner, send it, then evaluate once. If you must look early, treat anything short of the planned sample as a lean, which is exactly how the verdict describes it.

Test one variable at a time and start with the ones that move rates most. Subject lines change opens and, indirectly, replies. The opener and the core value sentence change replies the most. The call to action changes positive reply share more than total replies, so measure positive replies when testing it. Sending time, sender name and email length are worth testing only after the message itself is settled. Keep every other element, including the list, identical between arms.

Running an A/B test inside a sequence

In a multi-step sequence, randomize prospects into variants at the campaign level and keep them in the same variant for every step, so that a follow-up never mixes tones. Measure the metric at the same point for both arms, for example replies within 7 days of step one, and count the whole sequence's replies if the variable is the first email, since follow-ups inherit its context.

Split evenly. Unequal splits are valid but they slow the smaller arm and the test as a whole. Segment the list into halves that look alike on industry, title and company size, or randomize per prospect. When a test is significant, promote the winner, make the old winner the new control, and start the next test. ColdBox handles the split, keeps prospects locked to a variant across steps and reports replies per variant, so the numbers you paste here come straight from the campaign view.

FAQ

A/B Test Calculator questions

Straight answers, no fluff. Still stuck? Our deliverability team replies within a couple of hours.

Ask a human

It depends on your baseline rate and the smallest lift you care about. At a 3% reply rate, a 50% relative lift needs roughly 2,500 emails per variant at 95% confidence and 80% power, while a 20% lift needs about 14,000 per variant. Use the planner with your own baseline rather than a rule of thumb, since a 1% baseline needs about three times as many.

It is the probability of seeing a difference at least as large as yours if the two variants were actually identical. A p-value of 0.03 means a gap this size would appear by chance in about 3 tests out of 100 with no real difference. It is not the probability that variant B is better, which is a common misreading.

95% is the usual default and a fair balance for cold email. Use 90% when the change is cheap to roll back and you would rather move faster with more false positives. Use 99% when the decision is expensive, such as changing a sequence used by a whole team, and you can afford the larger sample it requires.

Ahead is not the same as proven. With low reply rates, a lead of a few replies is well within the range of random variation. The verdict compares your p-value to the threshold for the confidence level you chose. The how many more emails section estimates what it would take to confirm the lead if it is real.

You can, and the calculator supports it, but treat open rates with care. Privacy features in major mail clients prefetch images and inflate opens, which adds noise that has nothing to do with your subject line. Opens are fine for a quick directional read on subject lines; decide on replies or positive replies whenever you can.

Absolute difference is the gap in percentage points, for example 3.0% to 3.6% is 0.6 points. Relative uplift divides that gap by the baseline, so the same change is a 20% lift. Sample size planning uses relative lift because it scales sensibly across different baseline rates, while the confidence interval is shown in absolute points.

A significant result on a fair split is good evidence, but effects fade as lists, seasons and inboxes change. Promote the winner, keep it as the new control, and rerun on the next batch. If the second test agrees, the change is real. If it disagrees, you may have hit the one in twenty false positive that 95% confidence allows.

Start Free Today

Run every test on a clean split, automatically.

ColdBox randomizes prospects into variants, keeps them locked across every step and reports reply rates per variant, so you paste real numbers, not guesses. Free 7-day trial, no credit card.

Free trialNo credit cardSetup in 5 minutes