Reading
Stories Mode

Statistics in the Wild

~14 min read Lesson 1 of 3 in Module 12

Numbers Everywhere, Understanding Nowhere

Statistical claims flood the news every day. A headline announces that a new diet reduces heart disease risk by 40%. A study declares that people who sleep more than eight hours live longer. A politician asserts that crime has fallen 25% under their watch. Each of these claims might be true, misleading, or completely wrong — and the difference matters enormously for decisions that affect real lives.

The goal of this lesson is not to make you distrust all quantitative claims, but to give you the tools to evaluate them critically. After eleven modules of statistics, you are equipped to ask the right questions: What was the study design? Was there a comparison group? How large was the sample? What does the effect size actually mean in practical terms?

The Statistician's Reflex

Whenever you encounter a statistical claim, ask three questions before accepting it: Compared to what? (Is there a control group?) How big is the effect? (Is a statistically significant result also practically meaningful?) Who was measured? (Does the sample generalize to the population you care about?)

Common Mistakes in the Wild

Confusing correlation with causation is the most pervasive error. A headline declaring that coffee drinkers have lower rates of Alzheimer's disease does not mean coffee prevents Alzheimer's. Coffee drinkers may differ from non-drinkers in dozens of ways — diet, income, education, exercise — any of which could explain the association. Only a randomized experiment can establish causation; observational studies can only establish correlation.

Ignoring the base rate produces the base rate fallacy. Suppose a disease affects 1 in 1,000 people, and a test for it is 99% accurate (1% false positive rate). If you test positive, what is the probability you actually have the disease? Most people guess 99%. The correct answer is roughly 9%. Of every 1,000 people tested, about 10 false positives will occur (1% of the 999 disease-free people) while only 1 true positive is expected. The test is almost always wrong about a positive result, not because it is inaccurate, but because the disease is rare.

Survivorship bias occurs when we only see the cases that survived a selection process, making the survivors look systematically different from the original pool. The classic example: during World War II, analysts noted that returning aircraft had bullet holes concentrated in the fuselage and wings. They recommended reinforcing those areas — until Abraham Wald pointed out they should reinforce the areas with no holes, because planes hit there did not return. We only observed the survivors.

Absolute vs. Relative Risk

A drug that reduces heart attack risk from 2% to 1% has a relative risk reduction of 50% (sounds impressive) and an absolute risk reduction of 1 percentage point (modest). Always ask for both numbers. Relative risk reductions are routinely used in headlines because they sound larger.

Number Needed to Treat (NNT) = 1 / Absolute Risk Reduction

P-value misinterpretation remains epidemic even in scientific papers. A p-value of 0.03 does not mean there is a 3% chance the null hypothesis is true. It means: if the null hypothesis were true, there would be only a 3% chance of observing data at least as extreme as what we got. These are very different statements. A p-value tells you nothing about the probability that a hypothesis is true, nor about the size of the effect, nor about the practical importance of the finding.

Reproducibility and Open Science

The reproducibility crisis — also called the replication crisis — refers to the widespread finding, beginning around 2011, that a large fraction of published scientific results cannot be replicated when the experiments are repeated. A landmark 2015 study attempted to reproduce 100 psychology experiments and found that only about 39 showed results consistent with the original. Similar problems have been documented in medicine, economics, and neuroscience.

Several factors drive irreproducibility. Publication bias means journals preferentially publish positive results; null results go into the file drawer. P-hacking (also called “fishing”) refers to running many analyses and selectively reporting the one that crosses the significance threshold. HARKing (Hypothesizing After Results are Known) means generating a hypothesis to explain results after seeing the data, then presenting it as if it were pre-planned. Each of these practices inflates the number of false positives in the literature.

Open science is a set of practices designed to make research more transparent and reproducible. Its pillars include: sharing raw data and analysis code; pre-registering hypotheses and analysis plans before data collection; publishing null results; replicating key findings before treating them as established; and using pre-prints to accelerate access to findings. These are not just good habits — they are the foundation of a reliable scientific literature.

Pre-Registration

Pre-registration means publicly committing to your hypotheses, sample size, inclusion criteria, and analysis plan before you collect any data — and then following that plan. Registries like the Open Science Framework (OSF) and ClinicalTrials.gov provide time-stamped records that cannot be altered after data collection begins.

Pre-registration does not prevent exploratory analysis — it just requires you to label it honestly. A pre-registered confirmatory analysis is strong evidence; an exploratory analysis that emerged after looking at the data is hypothesis-generating, not hypothesis-confirming. The discipline of labeling each analysis correctly is one of the most powerful reforms in modern statistics.

Critics argue that pre-registration stifles creativity by forcing researchers to commit to a plan before understanding the data. The counter-argument is that it does not prevent exploration — it simply requires honest reporting. The confirmatory/exploratory distinction was always important; pre-registration makes it explicit and verifiable.

Meta-Analysis: Combining Studies

No single study is definitive. Every study has limited sample size, specific populations, particular protocols, and potential idiosyncratic flaws. Meta-analysis is the statistical synthesis of results across multiple independent studies on the same question. Done well, it produces the most reliable estimate of an effect that the existing literature can support.

The central idea of meta-analysis is to compute a weighted average of effect sizes across studies, giving more weight to studies with smaller standard errors (i.e., more precise estimates). Under the fixed-effects model, the pooled estimate is:

Meta-Analysis Pooled Effect (Fixed Effects)
\bar{\theta} = \frac{\displaystyle\sum_{i} w_i\,\hat{\theta}_i}{\displaystyle\sum_{i} w_i}, \quad w_i = \frac{1}{\mathrm{SE}_i^2}
Each study i contributes an effect size estimate θ̂i with weight wi = 1/SEi². Studies with smaller standard errors (larger samples, more precise measurement) receive proportionally more weight in the pooled estimate.

The fixed-effects model assumes all studies estimate the same true effect. The random-effects model relaxes this: it allows each study to estimate a slightly different true effect (due to variation in populations, protocols, or contexts), and the goal is to estimate the average effect across this distribution of true effects. Random-effects meta-analysis is more common in practice because effect sizes genuinely vary across study contexts.

Meta-analyses are visualized using forest plots: each study appears as a horizontal confidence interval on a vertical axis, with the size of the central square proportional to the study’s weight. The pooled estimate appears at the bottom as a diamond. A forest plot lets you see at a glance whether studies are consistent (the intervals overlap) or heterogeneous (they spread across very different values).

Meta-analyses can be misled by the same problems that corrupt primary studies. If positive results are more likely to be published, the pool of studies entering the meta-analysis is already biased. Funnel plots — which plot effect size against study precision — can reveal asymmetry that suggests publication bias: small, imprecise studies showing large effects on one side, with no counterbalancing studies on the other.

How to Read a Statistical Claim Critically

Armed with the concepts from this course, here is a practical checklist for evaluating any statistical claim you encounter — whether in a news headline, a scientific paper, or a policy report.

Critical Reading Checklist

1. What was the study design? — Observational? Randomized? Causal claims from observational studies warrant skepticism.

2. What is the comparison group? — Every claim of a difference requires a reference point. Compared to what?

3. What is the effect size? — Is the result reported in absolute terms, or only relative? Is it practically meaningful?

4. How large is the sample? — Small samples produce noisy estimates. Large samples can make trivial effects statistically significant.

5. Was the study pre-registered? — Pre-registered studies are less susceptible to p-hacking and HARKing.

6. Has it been replicated? — A single striking finding should be treated as preliminary until independently replicated.

7. Who funded it? — Industry-funded studies are more likely to report results favorable to the sponsor, not because of fraud, but through subtle choices in design and analysis.

Reading critically does not mean dismissing every statistical claim. It means asking these questions before updating your beliefs. When a study answers most of them well — randomized design, adequate power, pre-registered, independently replicated — it deserves substantial weight. When it answers few of them, it deserves proportionally less.

The next lesson surveys the frontier of statistics: survival analysis for time-to-event data, spatial and functional statistics, high-dimensional methods, and the evolving relationship between classical statistics and machine learning. These are the directions in which the field is currently moving.

Key Takeaways
  • Correlation is not causation. Observational studies establish associations; only randomized experiments establish causation. Always ask about confounders.
  • The base rate fallacy shows that even highly accurate tests produce mostly false positives when the condition is rare. Absolute risk and relative risk tell very different stories.
  • Survivorship bias arises whenever we only observe the units that made it through a selection process. The missing data is often the most informative data.
  • P-values do not measure the probability a hypothesis is true. They measure how surprising the data would be if the null were true. Statistical significance does not imply practical importance.
  • The reproducibility crisis reflects publication bias, p-hacking, and HARKing. Pre-registration, open data, and replication are the structural fixes.
  • Meta-analysis synthesizes evidence across studies using inverse-variance weighting. Funnel plots detect publication bias. Random-effects models accommodate genuine heterogeneity in effect sizes.
Previous Causal Inference Module Overview Next Lesson Advanced Topics Preview