Skip to main content

A/B Test Significance Calculator

Visitors and conversions for control and variant. The tool runs a two-proportion z-test and tells you whether to trust the lift.

Free toolsConversion & TestingReviewed September 2026

Control (A)

Variant (B)

Test settings

How sure you want to be before calling a winner. 95% is the common default.

Test type

Two-tailed tests for any difference. One-tailed tests only whether the variant is better.

Verdict

–

Enter visitors and conversions for both arms

Control rate
–
Variant rate
–
Relative lift
–
Variant vs control
Absolute lift
–
Percentage points
Z-score
–
P-value
–
Confidence (1 − p)
–
Not the chance B is better
Z needed
–
To reach your confidence level

Runs entirely in your browser. Nothing you enter is stored or sent anywhere. Last reviewed September 2026.

What significance does, and does not, tell you

A significant result means the gap you saw would be unusual if the two versions were really the same. That is all. The tool pools both arms into one conversion rate, works out how much two samples of this size would normally wander apart by chance (the standard error), and expresses your observed gap as a multiple of that wander (the z-score). A z-score near 2 or beyond is hard to get by luck alone, so the p-value drops below 0.05 and the result clears the 95% bar.

Significance is not the same as importance, and it is not a forecast. A 15% lift measured on 10,000 visitors per arm might really be anything from a few percent to well over 20%; the test tells you the direction is probably real, not that the number will hold. Small samples cut both ways: an underpowered test can miss a real winner and can also crown a fake one, which is why the tool warns when either arm has few conversions.

Two habits protect you. Decide the sample size before you start, using the sample size calculator, and do not stop early because the dashboard turned green. Checking repeatedly and stopping on a good day inflates false positives. And pick the test type before you look at the data. A one-tailed test asks only "is B better than A" and reaches significance more easily; a two-tailed test asks "is B different from A" and is the honest default when a variant could plausibly lose.

Formulas

p₁
= Control conversions ÷ Control visitors
p₂
= Variant conversions ÷ Variant visitors
Pooled p
= (Control conversions + Variant conversions) ÷ (Control visitors + Variant visitors)
SE
= √( Pooled p × (1 − Pooled p) × (1 ÷ n₁ + 1 ÷ n₂) )
z
= (p₂ − p₁) ÷ SE
p-value (two-tailed)
= 2 × (1 − Φ(|z|))
p-value (one-tailed, B better)
= 1 − Φ(z)
Relative lift
= (p₂ − p₁) ÷ p₁
Absolute lift
= p₂ − p₁
Significant when
p-value < 1 − Confidence level

Frequently asked questions

What does a p-value of 0.03 actually mean?

If the control and variant truly converted at the same rate, you would see a gap at least this large about 3% of the time purely from random variation in who happened to land in each group. It is not the probability that the variant is better, and it says nothing about the size of the real lift. A small p-value with a small sample can still come with a wide range of plausible true lifts, including some close to zero.

Should I use a one-tailed or two-tailed test?

Two-tailed is the safer default. It asks whether the variant is different in either direction, which is what you usually need to know, since a variant can lose. A one-tailed test asks only whether the variant is better, and it reaches significance with a smaller gap because all of the allowed error sits on one side. Use it only when you decided before the test that a loss and a tie would be treated identically, and say so when you report the result.

Can I stop the test as soon as it reaches 95%?

Not if you want the 95% to mean 95%. Checking daily and stopping on the first significant reading is called peeking, and it can push the real false-positive rate well above the 5% you think you are accepting, because a noisy test will cross the line at some point by chance. Decide the sample size in advance with the sample size calculator, run for whole weeks so weekday and weekend traffic are both included, and read the result once at the end.

More free tools

Browse all free tools →