About the A/B Test Significance Calculator
This A/B test significance calculator tells you whether the difference between two versions of a page, email or ad is real or just random noise. Enter the visitors and conversions for the control (A) and the variant (B), pick a confidence level, and it runs a two-proportion z-test to report the conversion rates, relative uplift, z-score, p-value and a clear significant / not significant verdict.
It is meant for marketers, product managers and CRO specialists deciding whether to ship a winning variant. Because stopping a test too early is the most common mistake, the calculator also estimates how many visitors per variant you would need to reliably detect the uplift you are seeing, at 80% statistical power.
The test assumes visitors are randomly split, each visitor converts at most once, and the sample is large enough for the normal approximation (at least a handful of conversions in each group).
How to use the a/b test significance calculator
- 1Enter visitors and conversions for the control (A).
- 2Enter visitors and conversions for the variant (B).
- 3Choose your confidence level — 95% is the common standard.
- 4Pick two-sided unless you decided in advance to test only for improvement.
- 5Read the verdict and p-value, and check the sample size needed before stopping the test.
Formula and method
The calculator uses a pooled two-proportion z-test. Each conversion rate is conversions ÷ visitors, and the pooled rate p̂ combines both groups under the assumption that there is no real difference. The z-score measures how many standard errors apart the two rates are, and the p-value is the probability of seeing a gap at least this large by chance. If the p-value is below 1 − confidence (0.05 at 95%), the result is significant.
The required sample size per variant uses the standard formula n = (z_α·√(2p̄(1−p̄)) + z_β·√(p_A(1−p_A) + p_B(1−p_B)))² ÷ (p_B − p_A)², with 80% power (z_β ≈ 0.842) and p̄ the average of the two rates.
- p_A, p_B
- Conversion rates of control and variant
- n_A, n_B
- Visitors in each group
- c_A, c_B
- Conversions in each group
- z_α
- Critical z for the confidence level (1.96 at 95% two-sided)
Worked examples
5.0% vs 6.0% on 5,000 visitors each
A converts 250/5,000 = 5% and B converts 300/5,000 = 6%, a 20% relative lift. The pooled z-test gives z ≈ 2.19 and p ≈ 0.028, below 0.05, so B wins at 95% confidence. To detect a lift this size with 80% power you would ideally have about 8,158 visitors per variant.
Small test that is not yet significant
B looks 27% better (3.81% vs 3.00%), but with only about 1,200 visitors per side the z-score is 1.09 and the p-value about 0.274 — far above 0.05. The gap could easily be chance, so keep the test running.
One-sided test at 90% confidence
A 6.0% vs 6.5% difference on 8,000 visitors each gives z ≈ 1.31. A one-sided p-value of about 0.096 is just below the 0.10 threshold for 90% confidence, so B is judged better — but this would not pass a two-sided 95% test.
Frequently asked questions
What does statistically significant mean in an A/B test?+
It means the difference you observed is unlikely to be due to random chance alone. At 95% confidence, a significant result has a p-value below 0.05: if the variants truly performed the same, a gap this large would appear less than 5% of the time.
How long should I run an A/B test?+
Run it until each variant reaches the required sample size and for at least one or two full weeks so weekday and weekend behaviour are both included. Decide the sample size before starting and avoid stopping as soon as the result looks significant.
Should I use a one-sided or two-sided test?+
Two-sided is the safer default because it detects both improvements and declines. Use one-sided only if you decided before the test that you care only whether B is better, since it makes significance easier to reach.
What is statistical power?+
Power is the probability that the test detects a real difference of a given size. 80% power is the common standard, meaning a true effect of that size would be missed 20% of the time. Low traffic tests are often underpowered.
Why is my big uplift not significant?+
With small samples, conversion rates swing a lot by chance, so even a 20–30% relative lift can fall within normal noise. Significance depends on both the size of the gap and the number of visitors and conversions behind it.