A/B Testing & Statistical Significance
Why 'B was higher' isn't proof of anything
Every A/B test compares a metric (usually conversion rate) between a control (A) and a variant (B) shown to different, randomly assigned slices of traffic. The problem is that random sampling always produces *some* difference between two groups even if the underlying true conversion rate is identical — flip a fair coin 100 times in two batches and the two batches won't show exactly 50 heads each. Statistical significance testing exists to answer one narrow question: is the observed difference large enough, given the sample size, that it's unlikely to be pure noise?
The null hypothesis and the p-value
The null hypothesis (H0) is the default assumption that there is no real difference between A and B — that any observed gap is due to random sampling variation alone. The test doesn't try to prove your variant works; it tries to see how surprising the observed data would be *if the null hypothesis were true*.
The p-value is the probability of seeing a difference this large (or larger) purely by chance, assuming H0 is true. A p-value of 0.03 means: if there were truly no difference between A and B, you'd still see a gap this big about 3% of the time from randomness alone. The conventional threshold is p < 0.05 (a 95% confidence level) — below that, you 'reject the null hypothesis' and call the result statistically significant. A p-value is not the probability that your variant is better, and it is not the magnitude of the improvement — both are common misreadings.
Statistical significance ≠ business significance
Statistical power and sample size
Statistical power is the probability that your test correctly detects a real effect when one actually exists — conventionally targeted at 80%. Power depends on three things you control before running the test: the baseline conversion rate, the minimum detectable effect (MDE) you care about catching, and the sample size per variant. Smaller effects and lower baseline rates both require dramatically larger samples to detect reliably — this is why a low-traffic checkout page test to detect a 2% relative lift can take months, while a high-traffic top-of-funnel test can reach significance in days.
The practical workflow: decide your baseline rate and the smallest lift worth caring about *before* the test starts, calculate the required sample size, and commit to running until that sample is reached — not stopping early because a p-value happened to dip below 0.05 partway through.
import scipy.stats as stats
import math
def required_sample_size(baseline_rate, mde_relative, alpha=0.05, power=0.80):
p1 = baseline_rate
p2 = baseline_rate * (1 + mde_relative)
z_alpha = stats.norm.ppf(1 - alpha / 2)
z_power = stats.norm.ppf(power)
pooled = (p1 + p2) / 2
numerator = (z_alpha * math.sqrt(2 * pooled * (1 - pooled)) +
z_power * math.sqrt(p1 * (1 - p1) + p2 * (1 - p2))) ** 2
denominator = (p2 - p1) ** 2
return math.ceil(numerator / denominator)
# Baseline 3% conversion rate, want to detect a 10% relative lift
n_per_variant = required_sample_size(0.03, 0.10)
print(n_per_variant) # ~ 51,000 visitors per variantSample size scales sharply with baseline rate and MDE
| Baseline conversion rate | Minimum detectable effect (relative) | Approx. sample size per variant |
|---|---|---|
| 10% | 20% relative lift | ~3,600 |
| 10% | 10% relative lift | ~14,500 |
| 3% | 20% relative lift | ~13,000 |
| 3% | 10% relative lift | ~51,000 |
| 1% | 10% relative lift | ~157,000 |
The pitfalls that invalidate most 'wins'
Two mistakes account for the vast majority of false positive test results reported in practice:
Peeking — checking the p-value repeatedly throughout the test and stopping as soon as it crosses 0.05. Because random noise fluctuates the p-value above and below 0.05 continuously during a test's run, checking daily and stopping at the first significant reading inflates your real false-positive rate far above the nominal 5% — in some simulations, well past 20-30%. The fix is either a fixed pre-calculated sample size with no early stopping, or a formal sequential testing method (like an alpha-spending function) explicitly designed to allow valid early stopping.
Multiple comparisons — running many simultaneous tests (or testing many metrics within one test) and treating any single significant result as a win. If you test 20 unrelated variants against a control, you'd expect roughly one to hit p < 0.05 by chance alone even if none actually work. This needs either a correction (like Bonferroni, dividing your significance threshold by the number of comparisons) or, more practically, pre-registering one primary metric per test.
Novelty effects and day-of-week bias
What's next
Sound experimentation depends on trustworthy underlying event data — a test result is only as good as the conversion event it's measuring. Reviewing how events are structured and named ties directly back into getting reliable test inputs.
Next: Event Taxonomy & Naming Conventions →
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.