A/B Testing Fundamentals
A/B testing (also called split testing or randomized controlled experimentation) is the gold standard for making data-driven decisions. It involves randomly assigning users to two or more variants to measure the causal impact of a change.
Key Concepts
| Term | Definition |
|---|---|
| Control (A) | The baseline/original version |
| Treatment (B) | The modified version being tested |
| Unit of Analysis | User, session, pageview, etc. |
| Randomization | Random assignment to eliminate confounders |
| Sample Size | Number of observations needed for reliable results |
| Statistical Significance | Probability that the observed effect is real |
| Effect Size | Magnitude of the difference between variants |
Statistical Foundations
Hypothesis Testing Framework
Z-Test for Proportions
T-Test for Continuous Metrics
Sample Size Calculation
For Conversion Rate Tests
For Continuous Metrics
Power Analysis Visualization
Sample Size Calculator
Statistical Significance
P-Value Interpretation
| P-Value | Interpretation | Action |
|---|---|---|
| < 0.01 | Very strong evidence against H₀ | Ship with confidence |
| 0.01 - 0.05 | Strong evidence against H₀ | Likely ship |
| 0.05 - 0.10 | Weak evidence | Consider more data |
| > 0.10 | No significant evidence | Don't ship |
Confidence Intervals
Effect Size Measures
| Effect Size | Cohen's d | Cohen's h |
|---|---|---|
| Small | 0.2 | 0.05 |
| Medium | 0.5 | 0.15 |
| Large | 0.8 | 0.25 |
Complete A/B Test Implementation
Common Pitfalls
1. Peeking Problem (Optional Stopping)
2. Multiple Comparisons
| Comparisons (k) | Family-wise Error Rate |
|---|---|
| 2 | 9.75% |
| 5 | 22.6% |
| 10 | 40.1% |
| 20 | 64.2% |
3. Novelty and Primacy Effects
4. Simpson's Paradox
Multi-Armed Bandits
Epsilon-Greedy Strategy
Upper Confidence Bound (UCB)
Thompson Sampling
Multi-Armed Bandit Implementation
Experimentation Platform Architecture
Key Takeaways
Practice Exercises
- Calculate the required sample size for baseline conversion 5%, MDE 10% relative lift, α=0.05, power=80%
- Implement a Thompson Sampling algorithm and compare against epsilon-greedy
- Research and implement a sequential testing framework
- When would you choose a multi-armed bandit over a traditional A/B test?
- How do you handle experiments where the metric takes weeks to mature?