A/B test sample size calculator
How many visitors each variant needs for a two-sided test of two proportions, how long that takes at your traffic, and whether the test is feasible at all.
What the result means
The sample is the number of visitors each arm needs before the test can tell a real effect of the stated size from chance, at the reliability you chose. The significance level is the share of the time a test would report a difference when there is none: at five percent two-sided, one test in twenty on an unchanged page reads as a change, in either direction. The power is the share of the time a real effect of the stated size is detected: at eighty percent, one real effect in five is missed.
A smaller effect needs more people because the noise in a rate does not shrink with the effect; only the sample shrinks it, and the effect is squared in the arithmetic, so half the effect needs four times the visitors. If you stop early, the moment the result looks good, the significance level you chose no longer holds: repeated looks raise the false positive rate well above it, and the result is a coin toss with a chart. Decide the sample in advance and read the test once, when it is reached.
Assumptions
Visitors are independent of each other, and each is assigned to one arm once. The metric is binary: a visitor converts or does not. The test runs to a fixed sample and is read once at the end, with no peeking. The variant converts at exactly the stated rate if the effect is real. Traffic is split as stated and the split holds.
Limitations
A sample ratio mismatch, a split that drifts from the one you set, invalidates the test whatever this says. Novelty and seasonality change the rates over the run; run a test for whole weekly cycles. A second metric on the same test, or a second variant, changes the arithmetic. Sequential designs and variance reduction methods such as CUPED change it too, and this tool does not model them. It is a fixed-horizon, two-arm, binary-metric calculation and nothing else.
Calculators differ in the variance they assume when there is no effect. This one uses the pooled rate of both arms, the textbook two-sample form; Evan Miller's calculator uses the baseline rate alone and gives about three percent fewer visitors on the reference case, 30,244 against 31,234. Neither is wrong. Both round the same decision, and the difference is smaller than the error in any baseline read from a single week.
Give us one funnel.
Twenty five minutes, about how you run experiments today. No access, no commitment.