Why 50 emails per variant tells you nothing
Cold email reply rates are small, commonly 1% to 5%, and that is what makes testing hard. If variant A gets 2 replies from 50 emails and variant B gets 4, B looks twice as good. In reality that is a difference of two replies, and the calculator will show a p-value well above 0.05: a gap that size or bigger shows up in a large share of tests where both variants are identical. The lower the base rate, the more emails it takes for a real difference to separate from noise.
Use the sample size planner before you start. At a 3% baseline reply rate, detecting a 20% relative lift (3% to 3.6%) at 95% confidence and 80% power needs about 14,000 emails per variant. Detecting a 50% lift (3% to 4.5%) needs roughly 2,500 per variant, and doubling the rate (3% to 6%) needs about 750. That is why most cold email tests should look for big swings, not polish.
The statistics behind the verdict
The calculator runs a two-proportion z-test. Each variant's rate is outcomes divided by emails sent. Under the null hypothesis that both variants share one true rate, it pools the two samples, computes the standard error of the difference, and expresses the observed difference as a z statistic. The two-sided p-value is the probability of a difference at least that large in either direction if nothing were different. The standard normal CDF is evaluated in the browser with the Abramowitz and Stegun approximation, accurate to about seven decimal places.
- Significant means p-value below alpha, where alpha is 0.10, 0.05 or 0.01 for 90, 95 or 99% confidence.
- The confidence interval for the difference uses each variant's own variance. If it excludes zero, the test is significant at that level.
- Minimum detectable effect (MDE) is the smallest relative lift your sample can reliably detect. Smaller MDE means far more emails, since sample size scales with the inverse square of the difference.
- 80% power means that if the true lift equals the MDE, four tests in five will reach significance. The remaining fifth will miss a real winner.
- The tool warns when a variant has fewer than 5 expected outcomes, where the normal approximation gets rough and the p-value should be read as a hint, not a proof.
The peeking problem and what to test
Checking results every day and stopping the moment p dips under 0.05 is the most common way cold email tests go wrong. Random noise crosses that line often on its way to nowhere, so a team that peeks ten times has a much higher false positive rate than the 5% it thinks it has. Fix the sample size with the planner, send it, then evaluate once. If you must look early, treat anything short of the planned sample as a lean, which is exactly how the verdict describes it.
Test one variable at a time and start with the ones that move rates most. Subject lines change opens and, indirectly, replies. The opener and the core value sentence change replies the most. The call to action changes positive reply share more than total replies, so measure positive replies when testing it. Sending time, sender name and email length are worth testing only after the message itself is settled. Keep every other element, including the list, identical between arms.
Running an A/B test inside a sequence
In a multi-step sequence, randomize prospects into variants at the campaign level and keep them in the same variant for every step, so that a follow-up never mixes tones. Measure the metric at the same point for both arms, for example replies within 7 days of step one, and count the whole sequence's replies if the variable is the first email, since follow-ups inherit its context.
Split evenly. Unequal splits are valid but they slow the smaller arm and the test as a whole. Segment the list into halves that look alike on industry, title and company size, or randomize per prospect. When a test is significant, promote the winner, make the old winner the new control, and start the next test. ColdBox handles the split, keeps prospects locked to a variant across steps and reports replies per variant, so the numbers you paste here come straight from the campaign view.