STAT 101
M06 · L04
Module 6 · Lesson 4

Common Pitfalls

P-hacking, multiple comparisons, and confusing statistical with practical significance — the habits that quietly broke science and how to avoid them.

01 / 12
STAT 101
M06 · L04
Fishing for significance

P-Hacking

  • Collect data until p < 0.05, then stop
  • Test many outcomes, report only the one that worked
  • Try different exclusion criteria until significance appears
  • Transform variables until the p-value improves
  • Choose one-tailed vs. two-tailed after seeing the data
02 / 12
STAT 101
M06 · L04
The real error rate

Familywise Error

FWER
\text{FWER} = 1 - (1 - \alpha)^m
20 tests at α=0.05
64% chance
of false positive
guaranteed
03 / 12
STAT 101
M06 · L04
Controlling FWER

Bonferroni Correction

Adjusted threshold
\alpha^* = \dfrac{\alpha}{m}

Divide α by the number of tests m. Simple, conservative. With m = 100 and α = 0.05, each test needs p ≤ 0.0005.

04 / 12
STAT 101
M06 · L04
Controlling FDR

Benjamini-Hochberg

BH Rule
p_{(k)} \le \dfrac{k}{m} \cdot q

Rank p-values, find largest k with p(k) ≤ (k/m)·q. Reject H₀ for all ranks 1…k. Controls the expected fraction of false discoveries. More powerful than Bonferroni.

05 / 12
STAT 101
M06 · L04
Statistical ≠ Important

Practical Significance

The trap
A drug lowers blood pressure by 0.1 mmHg — p < 0.0001 with n = 500,000. Statistically significant. Clinically meaningless.

P-value answers: “is this consistent with noise?” It does not answer: “does this matter?”

06 / 12
STAT 101
M06 · L04
Measuring what matters

Cohen’s d

Effect Size
d = \dfrac{\mu_1 - \mu_2}{\sigma_{\text{pooled}}}
  • |d| ≈ 0.2 — small effect
  • |d| ≈ 0.5 — medium effect
  • |d| ≈ 0.8 — large effect
07 / 12
STAT 101
M06 · L04
n → ∞ problem

Large Samples, Tiny p

Z-test statistic
z = \dfrac{\bar{x} - \mu_0}{\sigma / \sqrt{n}}

Doubling n multiplies z by √2. With huge n, any deviation from H₀ — no matter how trivial — becomes detectable. Effect size is sample-size-independent; p-value is not.

08 / 12
STAT 101
M06 · L04
The crisis

Reproducibility Crisis

Original studies significant
97%
Replications significant
36%

Open Science Collaboration (2015): 100 psychology studies replicated. Causes: p-hacking, publication bias, small n, lack of transparency.

09 / 12
STAT 101
M06 · L04
The solution

Pre-Registration

  • Commit hypotheses before data collection
  • Pre-specify outcomes, sample size, analysis plan
  • Eliminates researcher degrees of freedom
  • Distinguish confirmatory from exploratory analyses
  • Platforms: OSF, AsPredicted, AEA RCT Registry
10 / 12
STAT 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

11 / 12
STAT 101
Summary
Recap

What you learned

  • P-hacking exploits flexibility to find spurious significance
  • FWER = 1 − (1−α)ᵐ — grows fast with the number of tests
  • Bonferroni controls FWER: test at α/m
  • Benjamini-Hochberg controls FDR: more power for large test counts
  • Always report effect size (Cohen’s d) alongside p-values
  • Large n → tiny p, even for meaningless effects
  • Pre-register to make inference honest and reproducible
Module Complete
Hypothesis Testing — Module 6 ✓
12 / 12