Common Pitfalls
P-hacking, multiple comparisons, and confusing statistical with practical significance — the habits that quietly broke science and how to avoid them.
P-Hacking
- Collect data until p < 0.05, then stop
- Test many outcomes, report only the one that worked
- Try different exclusion criteria until significance appears
- Transform variables until the p-value improves
- Choose one-tailed vs. two-tailed after seeing the data
Familywise Error
Bonferroni Correction
Divide α by the number of tests m. Simple, conservative. With m = 100 and α = 0.05, each test needs p ≤ 0.0005.
Benjamini-Hochberg
Rank p-values, find largest k with p(k) ≤ (k/m)·q. Reject H₀ for all ranks 1…k. Controls the expected fraction of false discoveries. More powerful than Bonferroni.
Practical Significance
P-value answers: “is this consistent with noise?” It does not answer: “does this matter?”
Cohen’s d
- |d| ≈ 0.2 — small effect
- |d| ≈ 0.5 — medium effect
- |d| ≈ 0.8 — large effect
Large Samples, Tiny p
Doubling n multiplies z by √2. With huge n, any deviation from H₀ — no matter how trivial — becomes detectable. Effect size is sample-size-independent; p-value is not.
Reproducibility Crisis
Open Science Collaboration (2015): 100 psychology studies replicated. Causes: p-hacking, publication bias, small n, lack of transparency.
Pre-Registration
- Commit hypotheses before data collection
- Pre-specify outcomes, sample size, analysis plan
- Eliminates researcher degrees of freedom
- Distinguish confirmatory from exploratory analyses
- Platforms: OSF, AsPredicted, AEA RCT Registry
What you learned
- P-hacking exploits flexibility to find spurious significance
- FWER = 1 − (1−α)ᵐ — grows fast with the number of tests
- Bonferroni controls FWER: test at α/m
- Benjamini-Hochberg controls FDR: more power for large test counts
- Always report effect size (Cohen’s d) alongside p-values
- Large n → tiny p, even for meaningless effects
- Pre-register to make inference honest and reproducible