🎉 75% of content is free forever — Unlock Premium from $10/mo →
CW
💼 Servicesℹ️ About✉️ ContactView Pricing Plansfrom $10

A/B Testing and Experimentation

Module 4: Statistics and Probability🟢 Free Lesson

Advertisement

Control vs TreatmentControl (A)Original versionConversion: 10.0%n = 5,000Treatment (B)New versionConversion: 11.5%n = 5,000?Is B significantly better than A?

A/B Testing Fundamentals

A/B testing (also called split testing or randomized controlled experimentation) is the gold standard for making data-driven decisions. It involves randomly assigning users to two or more variants to measure the causal impact of a change.

A/B Testing Traffic FlowAll Traffic50%50%Control (A)Original VersionConversion: 10.0%Treatment (B)New VersionConversion: 11.5%Collect Metrics (n=5,000)Collect Metrics (n=5,000)Statistical Analysis (z-test / t-test)

Key Concepts

TermDefinition
Control (A)The baseline/original version
Treatment (B)The modified version being tested
Unit of AnalysisUser, session, pageview, etc.
RandomizationRandom assignment to eliminate confounders
Sample SizeNumber of observations needed for reliable results
Statistical SignificanceProbability that the observed effect is real
Effect SizeMagnitude of the difference between variants

Statistical Foundations

Hypothesis Testing Framework

Z-Test for Proportions

T-Test for Continuous Metrics

Sample Size Calculation

For Conversion Rate Tests

For Continuous Metrics

Power Analysis Visualization

Power vs Sample Size0.00.20.40.60.81.0Power0500100015002000250030003500Sample Size →Target 0.8n ≈ 2000Diminishing ReturnsBeyond n=2000, eachadditional sample addsless power gainPower curveTarget (0.8)Min sample size

Sample Size Calculator

Statistical Significance

P-Value Interpretation

P-ValueInterpretationAction
< 0.01Very strong evidence against H₀Ship with confidence
0.01 - 0.05Strong evidence against H₀Likely ship
0.05 - 0.10Weak evidenceConsider more data
> 0.10No significant evidenceDon't ship

Confidence Intervals

Effect Size Measures

Effect SizeCohen's dCohen's h
Small0.20.05
Medium0.50.15
Large0.80.25

Complete A/B Test Implementation

Common Pitfalls

1. Peeking Problem (Optional Stopping)

Peeking Problem (Optional Stopping)Sample Size →Test Statisticα = 0.05Stop early?"Looks good!"But is it?True path⚠ Inflates false positive rate from 5% to 30%+

2. Multiple Comparisons

Comparisons (k)Family-wise Error Rate
29.75%
522.6%
1040.1%
2064.2%

3. Novelty and Primacy Effects

Novelty and Primacy EffectsTime →Effect SizeLaunchNovelty EffectNew wears off quicklyInflated initial liftTrue EffectStable long-term impactPeak⚠ Early winners may be novelty effects — wait for stable signal

4. Simpson's Paradox

Simpson's ParadoxOverall: A beats B (5% vs 4%)MobileA: 2%B: 3%B wins!DesktopA: 8%B: 7%A wins!

Multi-Armed Bandits

Epsilon-Greedy Strategy

Upper Confidence Bound (UCB)

Thompson Sampling

Multi-Armed Bandit Implementation

Experimentation Platform Architecture

ConfigurationAllocation | --...

Key Takeaways

Practice Exercises

  1. Calculate the required sample size for baseline conversion 5%, MDE 10% relative lift, α=0.05, power=80%
  2. Implement a Thompson Sampling algorithm and compare against epsilon-greedy
  3. Research and implement a sequential testing framework
  4. When would you choose a multi-armed bandit over a traditional A/B test?
  5. How do you handle experiments where the metric takes weeks to mature?

Need Expert Data Science Help?

Get personalized tutoring, project support, or professional consulting.

Advertisement