This A/B testing sandbox was generated by Claude (Anthropic) for the STAT 101 course materials. Every number on it is computed in your browser and recomputed on every knob change: the pooled-proportion z statistic, both the two-sided and one-sided p-values, the pooled and unpooled standard errors, the lift and its confidence interval, the analytic power, the required per-arm sample size, and the whole power-vs-n curve. The false-positive and peeking rates come from a real seeded Monte-Carlo simulation (mulberry32), so the same seed always gives the same estimate. Nothing on this page is sketched.
The test and the interval use different standard errors on purpose. The z-test pools the two proportions — because under the null hypothesis they are equal — while the confidence interval for the lift does not pool, because it does not assume they are equal. Swapping the two is a classic mistake, so the page keeps them separate and marks which is which. At zero true effect the power equals α exactly: that number is the false-positive rate.
The two-proportion z-test, the difference-of-proportions interval, the normal-approximation power and sample-size formulas, and the optional-stopping (peeking) inflation are standard results. The conversion rates, sample size, alpha, target power and seed are illustrative inputs you set; the arithmetic done on them is exact. normCdf: Abramowitz & Stegun 26.2.17. normInv: Acklam (public domain). Seeded generator: mulberry32 (public domain). No number on this page came from a real experiment.
Course demo — reached from the Labs & Demos hub; the page itself is English‑only for now. Built for STAT-101 Module 11. Set a control rate p_A, a variant rate p_B and a per-arm sample size, and read the two-proportion z, the p-value and the confidence interval for the lift. Then see the power to detect that lift, the sample size a target power needs, and why peeking at the results early makes a 5% test reject far more than 5% of the time.