Reading
Stories Mode

A/B Testing

~14 min read Lesson 2 of 3 in Module 11

Experimentation at Scale

Every time you load a webpage, click a button, or see a headline, there is a good chance you are part of an experiment. Technology companies run thousands of A/B tests simultaneously — comparing version A (the control) against version B (the variant) to determine which produces better outcomes. Netflix tests thumbnail images. Amazon tests button colors. Google tests the shade of blue on its search links. This is experimental design as practiced at internet scale.

An A/B test is simply a randomized controlled experiment applied to a digital product. The principles from the previous lesson — randomization, control groups, power analysis — all apply here, but with a particular set of conventions, metrics, and failure modes specific to the online context.

The Core Idea

Randomly split users into two groups. Show group A the current experience. Show group B a modified experience. Measure whether the difference in outcomes is large enough to be explained by more than chance alone.

A/B test = randomized controlled experiment applied to product decisions

Setting Up: Hypothesis and Metric

Every A/B test begins with a well-defined hypothesis. A vague hypothesis like “the new design will be better” is not testable. A precise hypothesis specifies: what you are changing, what outcome you are measuring, and in what direction you expect it to move. For example: “Moving the call-to-action button above the fold will increase the click-through rate.”

The choice of primary metric is critical. You can measure almost anything — clicks, purchases, session length, return visits, revenue per user — but you must commit to a single primary metric before the test runs. Running the test and then choosing whichever metric moved is a form of p-hacking that inflates your false-positive rate.

Good metrics are sensitive (they change meaningfully when the treatment works), timely (they can be measured within the test window), and aligned with business goals. A metric that is easy to move but disconnected from real value is a vanity metric — it can lead you to ship changes that hurt in the long run.

Sample Size: Knowing When You Have Enough

Before launching an A/B test, you must determine how many users each group needs. This requires the same power analysis from the previous lesson, adapted to the metrics you are using.

For a binary outcome (did the user click or not?), the relevant quantities are: the baseline conversion rate p, the minimum detectable effect Δ (the smallest improvement worth caring about), the significance level α (typically 0.05), and the desired power 1−β (typically 0.80).

Sample Size for Proportions (Per Group)
n = \frac{(z_{\alpha/2} + z_{\beta})^2 \bigl[p_1(1-p_1) + p_2(1-p_2)\bigr]}{\Delta^2}
p̄ = (p₁ + p₂)/2 is the pooled proportion; zα/2 and zβ are critical values for significance and power; Δ = p₂ − p₁ is the minimum detectable effect.

A common mistake is setting the minimum detectable effect too small — trying to detect a 0.01% improvement when the baseline is 5%. This requires enormous sample sizes and long test durations. Instead, ask: what is the smallest effect that would actually influence our decision? Effects smaller than that are not worth detecting.

Running the Test: Randomization and Duration

Randomization unit matters. You can randomize at the level of individual page views, user sessions, or user identities. Randomizing by user identity (so each user always sees the same variant) is usually correct: it prevents the same user from experiencing both versions, which would contaminate the comparison.

The assignment mechanism must be truly random. A common approach is to hash the user ID and the experiment ID together, then take the result modulo 100 to assign bucket percentages. This ensures deterministic but effectively random assignment that is consistent across visits.

Test duration is set by your sample size calculation and the traffic rate, but there are two additional constraints. First, run the test for at least one full week to capture day-of-week effects — user behavior differs between weekdays and weekends. Second, run long enough to capture novelty effects: users sometimes respond differently to changes simply because they are new, and this effect fades after a few days.

Duration Checklist

1. Long enough to collect the required sample size.

2. At least one full week (7 days) to account for day-of-week effects.

3. Long enough to observe the steady-state behavior after novelty effects fade.

4. Short enough that the business context does not change significantly during the test.

Analyzing Results: t-test and Chi-Squared

Once the test concludes, you compare the two groups. The choice of test depends on the metric type.

For continuous metrics (average revenue per user, session duration), use a two-sample t-test. This tests whether the means of the two groups differ more than would be expected by chance.

For binary metrics (conversion rate, click-through rate), use a two-proportion z-test or equivalently a chi-squared test of independence. The test statistic is:

Two-Proportion Z-Test
z = \frac{\hat{p}_1 - \hat{p}_2}{\sqrt{\bar{p}\,(1-\bar{p})\!\left(\frac{1}{n_1}+\frac{1}{n_2}\right)}}
p̂₁ and p̂₂ are the observed conversion rates in control and variant; p̄ is the pooled proportion; n₁ and n₂ are the group sizes. Reject H₀ when |z| > zα/2.

Report the result with a confidence interval for the difference, not just the p-value. A statistically significant result that says “the variant improved conversion rate by 0.001% (95% CI: 0.0002%–0.002%)” gives far more information than “p = 0.03.” The confidence interval tells you the plausible range of the true effect size, which is what actually determines whether to ship.

Bayesian A/B Testing

The frequentist approach described above has a counterpart: Bayesian A/B testing. Instead of computing a p-value and comparing it to a threshold, the Bayesian approach models the conversion rates as random variables with prior distributions, updates them with the observed data, and computes the posterior probability that variant B is better than control A.

The key advantage is interpretability. “There is a 97% probability that variant B has a higher conversion rate” is a statement most product managers can intuitively understand. By contrast, a frequentist p-value of 0.03 is routinely misinterpreted as this same statement — but it is not. It is the probability of observing data this extreme if the null hypothesis were true, which is a subtly but importantly different claim.

Bayesian tests also handle the continuous monitoring problem more gracefully. In the frequentist framework, peeking at results mid-experiment and stopping early inflates the false-positive rate. In the Bayesian framework, updating beliefs as data arrives is philosophically natural, though careful implementation still requires thought about stopping rules.

Prior Distribution
p \sim \mathrm{Beta}(\alpha,\,\beta)
Beta(α, β) prior for conversion rate p. α − 1 prior successes, β − 1 prior failures.
Posterior Distribution
p \mid \text{data} \sim \mathrm{Beta}(\alpha + s,\; \beta + n - s)
After observing s successes in n trials, the posterior is Beta(α + s, β + n − s).

Common Pitfalls

Peeking is the most common mistake. You launch a test, check the results daily, and stop as soon as you see p < 0.05. This is statistically invalid. Under the null hypothesis, the p-value will wander below 0.05 roughly 5% of the time by chance — but if you check repeatedly, you greatly increase the chance of catching it on a false positive. Simulations show that peeking daily for a two-week experiment can inflate the false-positive rate from 5% to over 25%.

Multiple metrics create a similar problem. If you test 20 metrics and declare victory whenever any one reaches p < 0.05, you expect one false positive by pure chance. Apply a correction like Bonferroni (divide α by the number of tests) or Benjamini-Hochberg when testing many metrics.

Network effects and interference arise when users influence each other. If variant B users refer their friends, and those friends might have been assigned to control, then the two groups are no longer independent. Standard A/B testing assumes no interference between units — this is called the Stable Unit Treatment Value Assumption (SUTVA). When SUTVA is violated, cluster randomization or holdout designs are needed.

Sample ratio mismatch (SRM) occurs when the observed group sizes differ significantly from what was planned. If you assigned 50-50 but ended up with 47-53, something went wrong in the randomization or logging pipeline. Always check for SRM before interpreting results — it is a sign of a bug that can make any result meaningless.

Sequential Testing: Safe Early Stopping

Sometimes you want to stop a test early — for example, if the variant is causing harm, or if the effect is so large it is already clearly significant. The solution is sequential testing, also called group sequential design.

A sequential test pre-specifies a schedule of interim analyses and uses adjusted significance thresholds at each look. The most common approach is the O’Brien-Fleming boundary: the threshold for stopping early is very conservative (very small α) at early looks, and relaxes to near the nominal α at the final planned look. This controls the overall false-positive rate at exactly α while allowing early stopping when evidence is overwhelming.

An alternative is the always-valid p-value from the sequential testing literature, which remains valid regardless of when you stop. These methods allow continuous monitoring without inflating false-positive rates — the best of both worlds.

A/B Testing Checklist

1. Pre-specify — hypothesis, primary metric, sample size, and analysis plan before launch.

2. Randomize by user — consistent assignment prevents contamination within a user's session.

3. Check SRM — verify group sizes match the intended split before trusting any result.

4. Run the full duration — do not stop early unless using a sequential test with pre-specified boundaries.

5. Report effect sizes and confidence intervals — not just “significant” or “not significant.”

6. Correct for multiple comparisons — if testing many metrics or many variants simultaneously.

The next lesson extends from controlled experiments to the harder problem of causal inference from observational data — what to do when you cannot randomize, and how techniques like matching, propensity scores, and difference-in-differences let us approximate experimental conclusions from the data we have.

Key Takeaways
  • A/B testing is randomized controlled experimentation applied to digital products. Randomize by user identity, not by page view, to ensure consistent treatment.
  • Pre-specify your primary metric and sample size before launch. The minimum detectable effect should reflect the smallest change that would actually influence your shipping decision.
  • The two-proportion z-test (or chi-squared test) is standard for binary metrics; the two-sample t-test applies to continuous metrics. Always report confidence intervals alongside p-values.
  • Bayesian A/B testing provides more intuitive statements about the probability that the variant is better, and handles continuous monitoring more naturally than frequentist methods.
  • Peeking inflates false-positive rates dramatically. Sequential testing methods (O’Brien-Fleming, always-valid p-values) allow planned interim analyses without inflating error rates.
  • Always check for sample ratio mismatch before interpreting results. SRM indicates a bug in randomization or logging that makes any conclusion unreliable.
Previous Designing Experiments Module Overview Next Lesson Causal Inference