Statistical Significance
Statistical significance measures whether a test result reflects a real difference or just random variation. It determines when you can trust a test result.
Statistical significance is a measure of how confident you can be that a test result reflects a real difference rather than random chance.
What Statistical Significance Means in Marketing
When you run an A/B test and version B shows a higher conversion rate than version A, there are two possible explanations. Either B is genuinely better, or you just happened to measure during a period where B got lucky. Statistical significance is the tool for distinguishing between those two explanations.
It works through probability. If the difference between A and B could easily have occurred by chance, the result is not significant and you cannot act on it. If the difference is large enough, or the sample big enough, that chance alone is an unlikely explanation, the result is significant and you have grounds to prefer B.
The standard threshold in marketing testing is 95% confidence, meaning the result you observed would happen by chance in fewer than 5% of identical experiments where there was no real difference. That threshold is a convention, not a principle. Some teams work at 90%. Some at 99%. The right level depends on what you are changing and what the cost of an error is.
How Statistical Significance Works
The three factors that determine significance:
- Effect size. The bigger the difference between A and B, the easier it is to detect. A 20% improvement in conversion rate reaches significance faster than a 1% improvement.
- Sample size. More observations reduce the influence of random variation. A test on 100 visitors can rarely achieve significance. A test on 10,000 has enough data to detect real differences.
- Baseline conversion rate. A page converting at 2% needs a much larger sample to detect a meaningful change than a page converting at 20%, because low-conversion events are rare and noisy.
Before you start a test, calculate the sample size you need to detect the minimum effect you care about. Running the test without this step leads to either stopping too early or running indefinitely.
Statistical Significance Example
An e-commerce brand tests two product page layouts. After three days, layout B shows 12% more add-to-carts. The result is at 73% confidence. The marketing team calls it a win and rolls out B site-wide. Conversion rate then drops below the original. The early result was noise. Had they waited for 95% confidence with the required sample size, they would have seen the effect disappear before making the change.
Why Statistical Significance Matters for Marketers
Testing without statistical rigour is not really testing. It is just making decisions and attributing them to data.
The two errors pull in opposite directions. Stopping too early means acting on noise. Running too long means missing genuine improvements or continuing to test something you already know the answer to. Knowing what significance means keeps you honest about which situation you are in.
Frequently Asked Questions
What does a 95% confidence level mean in an A/B test?
It means that if there were no real difference between variant A and variant B, you would see a result this extreme or more extreme only 5% of the time by chance. It does not mean you are 95% sure variant B is better. It means the result is unlikely enough to be pure noise that you can act on it. The threshold is a convention, not a law.
How long should you run an A/B test before checking for significance?
Long enough to capture at least one full business cycle, typically one week minimum to account for day-of-week effects, and until you reach your pre-determined sample size. Checking early and stopping when you see significance is a common error called p-hacking. The result will often reverse when the test runs to full sample.
Can a result be statistically significant but commercially meaningless?
Yes, frequently. A test with a very large sample can detect tiny differences, even ones that have no practical impact on revenue. A 0.2% improvement in click-through rate might be statistically significant on a million impressions but generate so little additional revenue that it is not worth implementing. Significance answers whether the effect is real. It does not answer whether it matters.