Statistics in the Wild
Numbers flood the news daily. Most people cannot evaluate them. After this course, you can — and you have a responsibility to.
Three Questions
Ask these before accepting any statistical claim:
- Compared to what? — Is there a control group?
- How big is the effect? — Significant ≠ important
- Who was measured? — Does the sample generalize?
Correlation ≠ Causation
Observational studies can never fully rule out confounders. The coffee-and-Alzheimer’s headline? Coffee drinkers may simply exercise more, sleep better, have higher incomes.
Base Rate Fallacy
A 99%-accurate test for a 1-in-1,000 disease: what is the chance a positive test is correct?
Survivorship Bias
We only see what survived a selection process. WWII bullet-hole analysis: planes with hits in the fuselage returned. Reinforce the areas with no holes — those planes never came back.
P-value Myths
A p-value of 0.03 does not mean there is a 3% chance the null is true. It means: if the null were true, 3% chance of data this extreme.
Absolute vs. Relative Risk
Risk drops from 2% to 1%: a 50% relative reduction and a 1 percentage-point absolute reduction. Headlines always use relative — it sounds bigger.
Reproducibility Crisis
A 2015 study reproduced 100 psychology experiments. Only 39% replicated. Medicine, economics, and neuroscience face similar problems.
- Publication bias — null results disappear
- P-hacking — test many, report the winner
- HARKing — hypothesize after results known
Pre-Registration
Publicly commit to your hypothesis, sample size, and analysis plan before collecting data. Time-stamp it. Then follow it.
Meta-Analysis
Weighted average of effect sizes across independent studies. Precision = weight. More precise studies count more.
The Checklist
- Study design? — Observational or randomized?
- Comparison group? — Compared to what?
- Effect size? — Absolute, not just relative
- Pre-registered? — Reduces p-hacking risk
- Replicated? — Single results are preliminary
What you learned
Correlation ≠ causation. Base rates matter. Survivorship bias hides the failures. P-values are widely misread. Pre-registration and meta-analysis are medicine for the reproducibility crisis.