Glossary · ConversionTOFU
Statistical significance
The short answer
Statistical significance indicates how unlikely it is that a difference between test variants happened by chance alone. A result that's "significant at 95%" means that if there were truly no difference, you'd see a gap this large less than 5% of the time.
It does not mean the winning variant is 95% likely to be better, and it says nothing about whether the difference is large enough to matter.
Why tests produce false winners
- Peeking — checking daily and stopping the first time significance appears
- Small samples — a handful of conversions can swing wildly
- Multiple metrics — test enough metrics and one will look significant by chance
- Short durations — missing weekday and weekend differences
Doing it properly
- Calculate the sample size before starting
- Pick one primary metric in advance
- Run full weeks
- Don't stop early just because a result looks good
- Consider practical significance — is the lift worth acting on?
Avoiding false winners
- Decide sample size and duration before starting
- Run tests for full weekly cycles
- Don't stop early because one version leads
- Test one main change at a time
- Check results by segment (mobile vs desktop)
- Confirm winners with follow-up tests for big decisions
See A/B testing.
Significance in practice
A test shows version B converting 10% better, but with few conversions the difference could be chance. Waiting until the planned sample size is reached gives a more reliable answer.
Common mistakes
- Peeking at results and stopping early
- Running many tests and ignoring false positives
- Confusing significance with business impact
Frequently asked questions
Is 95% confidence always required?
It's a convention, not a law. Higher-risk decisions deserve stricter thresholds; low-risk ones can accept more uncertainty.
What if we don't have enough traffic?
Test bigger changes or use research-led improvements instead.
What does 95% confidence mean?
Roughly, that a result as extreme would be unlikely if there were no real difference. It's a convention, not a guarantee.
What if traffic is too low for significance?
Test bigger changes, run tests longer, or use qualitative research and before/after analysis.
Can we stop a test early if it looks like a winner?
Stopping early increases the risk of false positives; plan sample sizes in advance.
What confidence level should tests use?
Many teams use 90–95%, but the right level depends on the decision's risk.
Related terms
Full glossaryTalk to a strategist
Thirty minutes. Your numbers. A straight answer.
Book a strategy call with the people who'd actually run your account. We'll tell you what we'd do — or that you don't need us yet.