STAT 101
M6 · Lab
Hands-on lab
Read power, not just a p-value

Run a one-sample t-test in the power sandbox, and predict — before you read it off the screen — the effect size d, the critical value, the power at two sample sizes, and the false-positive rate that makes p-hacking work.

Effect size and noncentrality
\delta = d\sqrt{n}, \quad d = \frac{\mu - \mu_0}{\sigma}

A small p-value only tells you the data were surprising under H₀. It does not tell you the effect is real, or large, or that your study could even have found it. Power — the chance of rejecting H₀ when the effect is real — is the number that says whether an experiment was worth running. And at a true effect of zero, the rejection rate is exactly α: run enough tests and one will look significant by luck, which is the whole engine of p-hacking.

1 / 9
STAT 101
M6 · Lab
Set it up
One t-test

Open the sandbox. Each step names the one knob to change; leave everything else at these values. The readouts and panels print every number you predict.

Set these values
Test statistic -> One-sample t Tails -> Two-tailed True effect -> 0.5 Population SD -> 1 n -> 20 alpha -> 0.05

Watch the effect-size d, the critical-value and the power readouts — and the shaded null and alternative curves that meet at the critical value.

2 / 9
STAT 101
M6 · Lab
Step 1 of 5
The effect size, d

With the true effect at μ − μ₀ = 0.5 and the population SD at σ = 1, Cohen's d = (μ − μ₀)/σ. Predict d, then read the effect-size readout.

Expected

The readout shows d = 0.50 — a “medium” effect by Cohen's benchmarks. Note it has no n in it: effect size is what you are trying to detect, independent of sample size.

3 / 9
STAT 101
M6 · Lab
Step 2 of 5
The critical value

Keep the one-sample t test, n = 20, α = 0.05, two-tailed. The critical value is the 0.975 quantile of t on n − 1 = 19 degrees of freedom. Predict it, then read the critical-value readout.

Expected

The critical value is 2.093 — a sample t beyond ±2.093 is what gives p < 0.05 and rejects H₀.

4 / 9
STAT 101
M6 · Lab
Step 3 of 5
How much power at n = 20?

Same knobs: t test, n = 20, α = 0.05, two-tailed, d = 0.5. Power is P(reject H₀ | the effect is real). Predict the power, then read the power readout.

Expected

Power reads 0.565 — about 56.5%. Even with a real medium effect, this study misses it roughly 44% of the time. A non-significant result here proves nothing.

5 / 9
STAT 101
M6 · Lab
Step 4 of 5
Raise n to 80, watch power climb

Drag n from 20 up to 80, leaving every other knob alone. Predict the new power, then read it — and compare it with the 0.565 you saw at n = 20.

Expected

Power jumps to 0.993 — about 99.3%. Quadrupling n (20 → 80) turned a near-coin-flip study into a near-certain one, because the noncentrality δ = d√n grows with √n.

6 / 9
STAT 101
M6 · Lab
Step 5 of 5
The p-hacking trap

Now set the true effect to μ − μ₀ = 0, so H₀ is TRUE. The power readout becomes the false-positive rate. Predict what fraction of tests still reject H₀, then read it.

Expected

The rejection rate is 0.05 — exactly α. At a true effect of zero, 1 test in 20 is “significant” by chance. Run 20 tests, keep only the significant one, and you have manufactured a false finding — that is p-hacking.

7 / 9
STAT 101
M6 · Lab
Your turn
Open the sandbox

Everything above is waiting in the sandbox. Move the true effect and n and watch the two curves narrow and separate, drag α to see the critical value and the shaded areas trade size, and press Run to count real rejections against the theory.

8 / 9
STAT 101
M6 · Lab
Wrap-up
What you did
  • Read the effect size d = 0.50, an n-free measure of the effect
  • Found the two-tailed critical value at 19 df: 2.093
  • Measured power at n = 20 — only 0.565 for a real medium effect
  • Watched power climb to 0.993 at n = 80, because δ = d√n
  • Saw the false-positive rate sit at α = 0.05 when H₀ is true — the engine of p-hacking
9 / 9