Statistical Significance Analysis for A/B Tests
Run an A/B test and see p < 0.05? Stop. If you stop the test at the first sign of significance, the probability of a false positive result rises to 26%. Typical scenario: a designer redesigned a button, the test showed conversion improvement in two days, but a week later the effect disappeared. Peeking is the most expensive mistake in split experiments. We have analyzed 50+ projects and guarantee that with our approach you will avoid this and other pitfalls.
Why Is Statistical Significance Critical for A/B Tests?
Statistical significance is the mathematical confirmation that the difference between variants is not due to chance. Without it, you risk implementing a change that actually degrades metrics. Or, conversely, reject a profitable improvement because of noise. We use two approaches: Frequentist and Bayesian. Each solves its own class of problems.
How Frequentist and Bayesian Approaches Help Avoid Errors?
P-value — the probability of observing an effect as extreme as the one obtained, under the null hypothesis. A threshold of 0.05 is standard, but it does not reflect the effect size. Confidence Level (usually 95%) means we are willing to be wrong 5% of the time. Statistical Power (80%) — the ability to detect a real effect. MDE — the minimum effect the test will catch given the sample size.
Z-test for Proportions
from scipy.stats import proportions_ztest, chi2_contingency
import numpy as np
def analyze_test(control_n, control_conv, variant_n, variant_conv, alpha=0.05):
cr_control = control_conv / control_n
cr_variant = variant_conv / variant_n
relative_lift = (cr_variant - cr_control) / cr_control * 100
# Z-test (applicable if n > 30)
counts = np.array([variant_conv, control_conv])
nobs = np.array([variant_n, control_n])
z_stat, p_value = proportions_ztest(counts, nobs, alternative='two-sided')
# Confidence interval for the difference
se = np.sqrt(
cr_control * (1 - cr_control) / control_n +
cr_variant * (1 - cr_variant) / variant_n
)
diff = cr_variant - cr_control
z_crit = 1.96 # for 95% CI
ci_low = diff - z_crit * se
ci_high = diff + z_crit * se
print(f"Control: {cr_control:.3%} ({control_conv}/{control_n})")
print(f"Variant: {cr_variant:.3%} ({variant_conv}/{variant_n})")
print(f"Lift: {relative_lift:+.1f}%")
print(f"95% CI: [{ci_low:.3%}, {ci_high:.3%}]")
print(f"P-value: {p_value:.4f}")
print(f"Significant: {'YES ✓' if p_value < alpha else 'NO ✗'}")
return p_value < alpha
analyze_test(
control_n=3842, control_conv=115,
variant_n=3891, variant_conv=148
)
Chi-square Test (Alternative to Z-test)
from scipy.stats import chi2_contingency
contingency = np.array([
[control_conv, control_n - control_conv], # Control: converts, not converts
[variant_conv, variant_n - variant_conv] # Variant: converts, not converts
])
chi2, p_value, dof, expected = chi2_contingency(contingency)
print(f"Chi2: {chi2:.4f}, p={p_value:.4f}")
Chi-square and Z-test give identical results for two groups.
What Is Peeking and How to Avoid It?
Peeking — stopping a test as soon as p < 0.05 appears, without waiting for the calculated sample size. This inflates the Type I error rate to 26% at alpha=0.05. Solution: pre-calculate the required sample size and do not interrupt the test until it is reached.
# Wrong: check every day and stop when p < 0.05
# Correct: calculate sample size in advance, stop only after reaching it
def required_sample_size(baseline_cr, mde, alpha=0.05, power=0.8):
from scipy import stats
import math
p1, p2 = baseline_cr, baseline_cr * (1 + mde)
p_avg = (p1 + p2) / 2
z_a = stats.norm.ppf(1 - alpha/2)
z_b = stats.norm.ppf(power)
n = ((z_a * math.sqrt(2 * p_avg * (1-p_avg)) +
z_b * math.sqrt(p1*(1-p1) + p2*(1-p2))) / (p2-p1)) ** 2
return math.ceil(n)
n = required_sample_size(baseline_cr=0.03, mde=0.15)
print(f"Run test until {n} users per variant reached")
For multiple comparisons, use Bonferroni correction:
# Bonferroni correction for multiple comparisons
n_comparisons = 4 # 4 variants vs control
corrected_alpha = 0.05 / n_comparisons # = 0.0125
# Or FDR (Benjamini-Hochberg)
from statsmodels.stats.multitest import multipletests
p_values = [0.03, 0.07, 0.01, 0.04]
reject, corrected_p, _, _ = multipletests(p_values, alpha=0.05, method='fdr_bh')
Bayesian A/B Analysis: Probabilistic Approach
An alternative to the frequentist approach — probability that a variant is better:
import numpy as np
def bayesian_ab_test(control_conv, control_n, variant_conv, variant_n, samples=100000):
"""Posterior distribution via Beta distribution"""
# Prior: Beta(1,1) = uniform distribution
control_posterior = np.random.beta(
control_conv + 1,
control_n - control_conv + 1,
samples
)
variant_posterior = np.random.beta(
variant_conv + 1,
variant_n - variant_conv + 1,
samples
)
prob_variant_better = (variant_posterior > control_posterior).mean()
expected_lift = (variant_posterior - control_posterior).mean() / control_posterior.mean() * 100
print(f"Probability variant is better: {prob_variant_better:.1%}")
print(f"Expected lift: {expected_lift:+.1f}%")
print(f"Credible interval: [{np.percentile(variant_posterior - control_posterior, 2.5):.3%}, "
f"{np.percentile(variant_posterior - control_posterior, 97.5):.3%}]")
bayesian_ab_test(115, 3842, 148, 3891)
Bayesian approach gives the probability that the variant is better, accelerating decision-making by 20% compared to Frequentist in multiple test scenarios.
Frequentist vs Bayesian: When to Use Which?
| Criterion | Frequentist | Bayesian |
|---|---|---|
| Interpretation | p-value, CI | Probability of hypothesis |
| Required sample size | Pre-fixed | Flexible, can monitor |
| Incorporates prior data | No | Yes (prior) |
| Computational complexity | Low | Higher (simulations) |
| Popularity | Classic, industry standard | Modern, intuitive |
| Situation | Decision |
|---|---|
| p < 0.05, lift > 0 | Launch variant |
| p > 0.05, low traffic | Continue test |
| p > 0.05, reached sample size | No significant effect, close test |
| p < 0.05, lift negative | Keep control |
| One segment significant, another not | Interaction analysis, segmented deployment |
Process and What’s Included
- Analytics — we analyze your current testing scheme, goals, and metrics.
- Design — choose the optimal method (Frequentist/Bayesian), calculate sample size.
- Implementation — integrate scripts or connect a library (e.g.,
scipy+statsmodels). - Testing — simulate on historical data, verify correctness.
- Deploy — set up an automated dashboard with results, documentation.
The deliverable includes the analysis source code (Python/R/JS) with comments, calculation of required sample size for your parameters, integration with your tracking system (Google Analytics, Mixpanel, custom logs), training for your team on result interpretation, and support for 2 weeks after deployment. We guarantee the correctness of calculations and accuracy of conclusions — our experience is confirmed by dozens of successful projects.
Timeframe and Cost
Setting up the statistical significance analysis with automatic sample size calculation and Bayesian/Frequentist choice takes 1–2 business days. The cost is calculated individually based on integration complexity — typically from $60 to $240. On average, clients reduce analysis time by 30% and avoid losses from incorrect decisions, which can cost a company up to $1,200 monthly. Schedule a consultation — contact us today!
Checklist of Common Mistakes
- Didn’t pre-calculate sample size.
- Stopped the test at the first p < 0.05.
- Forgot about multiple comparisons.
- Used p-value as the sole criterion without considering effect size.
- Didn’t segment the audience (e.g., different devices).
Contact us to set up reliable statistical analysis for your A/B tests and make confident decisions. Order a consultation on statistical significance calculation today — we’ll help you avoid mistakes and save your budget.







