STAT 101
M11 · L02
Module 11

A/B Testing

Randomized controlled experiments applied to digital products. How tech companies test thousands of changes at once — and how to do it rigorously.

01 / 13
STAT 101
M11 · L02
The Framework

Control vs. Variant

Split users randomly. Show A (control) to one group and B (variant) to the other. Measure whether the difference is larger than chance.

Core principle
A/B test = randomized controlled experiment applied to product decisions.
02 / 13
STAT 101
M11 · L02
Step 1

Hypothesis & Metric

Pre-specify what you are changing, what you are measuring, and the expected direction. Commit to one primary metric before launch.

  • Sensitive — changes when the treatment works
  • Timely — measurable within the test window
  • Aligned — connected to real business value
03 / 13
STAT 101
M11 · L02
Step 2

Sample Size

Set the minimum detectable effect before launch. Smaller effects need more users — aim for the smallest change that would actually influence your shipping decision.

Significance
α = 0.05
Power
80%
04 / 13
STAT 101
M11 · L02
Step 3

Randomize by User

Assign users to groups by hashing their ID — not by page view. This gives each user a consistent experience and prevents contamination within a single session.

Assignment rule
hash(user_id + experiment_id) mod 100 → bucket assignment
05 / 13
STAT 101
M11 · L02
Step 4

Test Duration

Run long enough for the required sample and to cover natural behavior cycles. Novelty effects fade after a few days — steady-state behavior is what matters.

Min 7 days — capture full week cycle
Novelty period — wait for behavior to stabilize
Fixed end date — set before launch, not after
06 / 13
STAT 101
M11 · L02
Step 5

Analyze the Results

Choose the test based on the metric type. Always report a confidence interval for the effect size — not just a p-value.

  • Continuous — two-sample t-test (mean revenue, session length)
  • Binary — two-proportion z-test (click rate, conversion)
  • Count — Poisson test or chi-squared
07 / 13
STAT 101
M11 · L02
Two-Proportion Test

The Test Statistic

For conversion rates: compare p̂₁ and p̂₂ relative to pooled variance.

Two-proportion z-test
z = \frac{\hat{p}_1 - \hat{p}_2}{\sqrt{\bar{p}\,(1-\bar{p})\!\left(\frac{1}{n_1}+\frac{1}{n_2}\right)}}
08 / 13
STAT 101
M11 · L02
Alternative Approach

Bayesian A/B Testing

Instead of a p-value, compute: “P(B > A)” — the probability the variant is better. More intuitive and handles continuous monitoring more naturally.

Prior
Beta(α, β)
Posterior
Beta(α+s, β+n−s)
09 / 13
STAT 101
M11 · L02
Pitfall #1

The Peeking Problem

Checking results daily and stopping when p < 0.05 inflates false-positive rates from 5% to over 25%. The p-value wanders — catching it low is not a real signal.

Solution
Pre-register your end date. Use sequential testing if early stopping is needed.
10 / 13
STAT 101
M11 · L02
Safe Early Stopping

Sequential Testing

Pre-specify interim analyses with adjusted thresholds. The O’Brien-Fleming boundary is conservative early and relaxes toward the final look — controlling overall error at exactly α.

  • O’Brien-Fleming — classic sequential boundary
  • Always-valid p-values — valid at any stopping time
  • Bayesian stopping rules — stop when P(B>A) exceeds threshold
11 / 13
STAT 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
STAT 101
Summary
Recap

What you learned

Pre-specify everything. Randomize by user. Check SRM. Run the full duration. Report confidence intervals. Avoid peeking — or use sequential testing if you must look early.

Next Lesson
13 / 13