Orodja/Testing

A/B Test Significance Calculator

Check whether your A/B test result is statistically significant. Enter visitors and conversions for both variants and get the p-value, uplift, and a plain verdict.

Variant A (control)
Variant B

Enter visitors and conversions for both variants to see the result.

What this calculator does

It runs a two-proportion z-test on your variants, the standard test for comparing two conversion rates. You get each variant's conversion rate, the relative uplift, the p-value, and whether the difference clears the usual 95% confidence bar. Everything runs in your browser.

How to read the result

The p-value is the probability of seeing a difference at least this large if the two variants actually performed the same. Below 0.05, the convention is to call the result statistically significant. That is a convention, not a law of nature: at p = 0.05, roughly one test in twenty will look like a winner by pure chance.

Significance also says nothing about size. A test can be significant with a tiny uplift that never pays back the engineering time, or inconclusive with a large uplift simply because the sample is small. Look at the uplift and its practical value, not just the verdict.

Common ways A/B tests go wrong

The classic mistake is peeking: checking the test daily and stopping the moment it shows significance. That inflates false positives badly, because you are giving randomness many chances to cross the line. Decide the sample size up front (the sample size calculator does this), run until you reach it, then read the result once.

Other traps: running many variants without correcting for multiple comparisons, testing during an unusual week, and declaring winners on segments you picked after seeing the data. When a result looks too good, it usually is. Track your conversions in a funnel so you can see where variants actually differ.

Pogosto zastavljena vprašanja

What does statistically significant mean?

It means the difference between variants is unlikely to be pure chance. At the usual 95% level, a significant result would occur less than 5% of the time if the variants truly performed the same.

What is a good p-value for an A/B test?

The convention is below 0.05, which corresponds to 95% confidence. For high-stakes changes, teams sometimes require 0.01. The lower the p-value, the less likely the result is noise.

How long should I run an A/B test?

Until you reach the sample size you calculated before starting, and ideally over at least one or two full weeks so weekday and weekend behavior are both represented. Stopping early because the result looks significant inflates false positives.

Can I test more than two variants?

Yes, but each extra comparison raises the chance that one looks significant by luck. Either correct for multiple comparisons (a Bonferroni correction is the simple option) or stick to fewer, better-reasoned variants.

My test is not significant. Is it a failure?

No. An inconclusive test tells you the change did not move the metric enough to detect at your sample size. That is useful information: ship whichever variant you prefer for other reasons, or test a bolder change.

Celotna zbirka orodij v enem zavihku

Analyse združuje analitiko brez piškotkov, lijake, zadrževanje in SEO motor, ki piše in objavlja raziskane objave. Preizkusi vse funkcije brezplačno 14 dni.

Začni brezplačno preizkusno obdobje