Run a one-sample t-test in the power sandbox, and predict — before you read it off the screen — the effect size d, the critical value, the power at two sample sizes, and the false-positive rate that makes p-hacking work.
A small p-value only tells you the data were surprising under H₀. It does not tell you the effect is real, or large, or that your study could even have found it. Power — the chance of rejecting H₀ when the effect is real — is the number that says whether an experiment was worth running. And at a true effect of zero, the rejection rate is exactly α: run enough tests and one will look significant by luck, which is the whole engine of p-hacking.
Open the sandbox. Each step names the one knob to change; leave everything else at these values. The readouts and panels print every number you predict.
Test statistic -> One-sample t Tails -> Two-tailed
True effect -> 0.5 Population SD -> 1 n -> 20 alpha -> 0.05
Watch the effect-size d, the critical-value and the power readouts — and the shaded null and alternative curves that meet at the critical value.
With the true effect at μ − μ₀ = 0.5 and the population SD at σ = 1, Cohen's d = (μ − μ₀)/σ. Predict d, then read the effect-size readout.
The readout shows d = 0.50 — a “medium” effect by Cohen's benchmarks. Note it has no n in it: effect size is what you are trying to detect, independent of sample size.
Keep the one-sample t test, n = 20, α = 0.05, two-tailed. The critical value is the 0.975 quantile of t on n − 1 = 19 degrees of freedom. Predict it, then read the critical-value readout.
The critical value is 2.093 — a sample t beyond ±2.093 is what gives p < 0.05 and rejects H₀.
Same knobs: t test, n = 20, α = 0.05, two-tailed, d = 0.5. Power is P(reject H₀ | the effect is real). Predict the power, then read the power readout.
Power reads 0.565 — about 56.5%. Even with a real medium effect, this study misses it roughly 44% of the time. A non-significant result here proves nothing.
Drag n from 20 up to 80, leaving every other knob alone. Predict the new power, then read it — and compare it with the 0.565 you saw at n = 20.
Power jumps to 0.993 — about 99.3%. Quadrupling n (20 → 80) turned a near-coin-flip study into a near-certain one, because the noncentrality δ = d√n grows with √n.
Now set the true effect to μ − μ₀ = 0, so H₀ is TRUE. The power readout becomes the false-positive rate. Predict what fraction of tests still reject H₀, then read it.
The rejection rate is 0.05 — exactly α. At a true effect of zero, 1 test in 20 is “significant” by chance. Run 20 tests, keep only the significant one, and you have manufactured a false finding — that is p-hacking.
Everything above is waiting in the sandbox. Move the true effect and n and watch the two curves narrow and separate, drag α to see the critical value and the shaded areas trade size, and press Run to count real rejections against the theory.
- Read the effect size d = 0.50, an n-free measure of the effect
- Found the two-tailed critical value at 19 df: 2.093
- Measured power at n = 20 — only 0.565 for a real medium effect
- Watched power climb to 0.993 at n = 80, because δ = d√n
- Saw the false-positive rate sit at α = 0.05 when H₀ is true — the engine of p-hacking