A/B Testing
Randomized controlled experiments applied to digital products. How tech companies test thousands of changes at once — and how to do it rigorously.
Control vs. Variant
Split users randomly. Show A (control) to one group and B (variant) to the other. Measure whether the difference is larger than chance.
Hypothesis & Metric
Pre-specify what you are changing, what you are measuring, and the expected direction. Commit to one primary metric before launch.
- Sensitive — changes when the treatment works
- Timely — measurable within the test window
- Aligned — connected to real business value
Sample Size
Set the minimum detectable effect before launch. Smaller effects need more users — aim for the smallest change that would actually influence your shipping decision.
Randomize by User
Assign users to groups by hashing their ID — not by page view. This gives each user a consistent experience and prevents contamination within a single session.
Test Duration
Run long enough for the required sample and to cover natural behavior cycles. Novelty effects fade after a few days — steady-state behavior is what matters.
Analyze the Results
Choose the test based on the metric type. Always report a confidence interval for the effect size — not just a p-value.
- Continuous — two-sample t-test (mean revenue, session length)
- Binary — two-proportion z-test (click rate, conversion)
- Count — Poisson test or chi-squared
The Test Statistic
For conversion rates: compare p̂₁ and p̂₂ relative to pooled variance.
Bayesian A/B Testing
Instead of a p-value, compute: “P(B > A)” — the probability the variant is better. More intuitive and handles continuous monitoring more naturally.
The Peeking Problem
Checking results daily and stopping when p < 0.05 inflates false-positive rates from 5% to over 25%. The p-value wanders — catching it low is not a real signal.
Sequential Testing
Pre-specify interim analyses with adjusted thresholds. The O’Brien-Fleming boundary is conservative early and relaxes toward the final look — controlling overall error at exactly α.
- O’Brien-Fleming — classic sequential boundary
- Always-valid p-values — valid at any stopping time
- Bayesian stopping rules — stop when P(B>A) exceeds threshold
What you learned
Pre-specify everything. Randomize by user. Check SRM. Run the full duration. Report confidence intervals. Avoid peeking — or use sequential testing if you must look early.