Was ist Statistical significance?
A threshold indicating a test result is unlikely to be random noise, commonly set at 95 percent confidence in A/B testing.
Statistical significance is the evidentiary bar for saying an A/B test result is unlikely to be luck. Conversion counts bounce around randomly, so a variant will often look ahead by chance; a significance test asks how probable a gap this large would be if the versions were actually identical. If that probability (the p-value) is below your threshold, conventionally 5 percent, the result is "significant at 95 percent confidence."
Two misreadings cause most damage. First, 95 percent confidence is not a 95 percent chance the variant is better; it means results this extreme would occur under 5 percent of the time by luck alone, which also implies that one in twenty null tests will "win" spuriously, more if you track many metrics or peek repeatedly. Second, significance is not size: with enough traffic a 0.1 percent lift will reach significance while being worth nothing, and an underpowered test can miss a real 20 percent lift. Always read significance together with the effect size and the sample size that produced it.
The discipline that keeps the math honest: fix the metric, threshold, and sample size in advance, and evaluate once at the end.
Verwandte Begriffe
Sieh dir diese Metriken auf deiner eigenen Website an
Analyse misst jede Metrik in diesem Glossar ohne Cookies und ohne Consent-Banner. Analytics, Funnels und eine SEO-Engine in einem Tab. 14 Tage kostenlos.
Kostenlos testen