Glossarju/Experimentation

X'inhu Statistical significance?

A threshold indicating a test result is unlikely to be random noise, commonly set at 95 percent confidence in A/B testing.

Statistical significance is the evidentiary bar for saying an A/B test result is unlikely to be luck. Conversion counts bounce around randomly, so a variant will often look ahead by chance; a significance test asks how probable a gap this large would be if the versions were actually identical. If that probability (the p-value) is below your threshold, conventionally 5 percent, the result is "significant at 95 percent confidence."

Two misreadings cause most damage. First, 95 percent confidence is not a 95 percent chance the variant is better; it means results this extreme would occur under 5 percent of the time by luck alone, which also implies that one in twenty null tests will "win" spuriously, more if you track many metrics or peek repeatedly. Second, significance is not size: with enough traffic a 0.1 percent lift will reach significance while being worth nothing, and an underpowered test can miss a real 20 percent lift. Always read significance together with the effect size and the sample size that produced it.

The discipline that keeps the math honest: fix the metric, threshold, and sample size in advance, and evaluate once at the end.

Termini relatati

Ara dawn il-metriċi fuq is-sit tiegħek stess

Analyse jsegwi kull metrika f'dan il-glossarju mingħajr cookies u mingħajr banner ta' kunsens. Analitiċi, funnels, u magna SEO f'tab wieħed. B'xejn għal 14-il jum.

Ibda prova b'xejn