The Seduction of Significant Results
Statistical hypothesis testing is one of science’s most powerful tools — and one of its most abused. A p-value below 0.05 has become a cultural threshold for “discovery,” and the pressure to clear that bar has quietly corrupted research practices across fields. In this lesson we examine the most dangerous pitfalls: p-hacking, the multiple comparisons problem, confusing statistical with practical significance, and the reproducibility crisis that these habits have produced.
Understanding these pitfalls is not merely academic. It will make you a more careful analyst, a more skeptical reader of published research, and a more honest communicator of your own results.
P-Hacking: Fishing for Stars
P-hacking (also called data dredging, data fishing, or researcher degrees of freedom) refers to the practice of trying many analyses on a dataset until one yields p < 0.05, and then reporting only that result. Because every test has a 5% chance of a false positive under H₀, running enough analyses virtually guarantees finding something “significant” — even in random noise.
P-hacking takes many forms: collecting data until the p-value crosses the threshold, testing multiple outcome variables and reporting only the one that worked, applying different exclusion criteria until significance appears, transforming variables (log, square root) until the p-value improves, or choosing between one-tailed and two-tailed tests after seeing the data. None of these practices is necessarily dishonest in isolation, but doing them without correction — and without disclosing them — inflates Type I error far above α.
Imagine testing whether jelly beans cause acne. You run 20 separate tests, one for each color. At α = 0.05 you expect one false positive by chance. One color (green) gives p = 0.04. The headline: “Green Jelly Beans Linked to Acne!” Without knowing the other 19 tests were run, the single published result looks compelling. This is the multiple comparisons problem in its purest form.
The Multiple Comparisons Problem
Whenever you perform m independent hypothesis tests simultaneously — or examine m outcomes, m subgroups, or m time points — the probability that at least one test yields a false positive is no longer α. Under the assumption that all null hypotheses are true, the familywise error rate (FWER) is:
This is not a flaw in p-values per se — it is a flaw in how they are used. A p-value is valid for a single pre-specified test. Applying it naively across many tests without adjustment is the error.
Corrections: Bonferroni and Benjamini-Hochberg
Two main correction strategies exist, each targeting a different error criterion.
Bonferroni correction controls the FWER. It simply divides α by the number of tests m, producing an adjusted threshold α* = α/m. Each individual test must clear this stricter bar to be declared significant:
When many hypotheses are tested simultaneously — for instance, tens of thousands of genes in a genomics study — Bonferroni is extremely conservative. A less stringent alternative controls the False Discovery Rate (FDR): the expected proportion of rejected hypotheses that are actually null.
The Benjamini-Hochberg (BH) procedure controls FDR at level q. Rank the m p-values from smallest to largest: p₁ ≤ p₂ ≤ … ≤ pᵐ. Find the largest k such that:
Bonferroni: when false positives are very costly and m is modest (e.g., clinical trials testing a handful of endpoints). Benjamini-Hochberg: when discoveries are many and some false positives are acceptable in exchange for higher power (e.g., genomics, neuroimaging, large-scale surveys). When in doubt: pre-specify which correction you will apply, and stick to it.
Statistical Significance vs. Practical Significance
Perhaps the most pervasive misunderstanding in statistics: a statistically significant result is not necessarily important. A p-value answers only one question: “Is the observed effect consistent with random chance under H₀?” It says nothing about whether the effect is large enough to matter in practice.
Effect size is the right measure for practical significance. For comparing two means, Cohen’s d is the most common effect size statistic:
A drug that lowers blood pressure by 0.1 mmHg with p < 0.0001 (in a study of 500,000 patients) is statistically significant but clinically meaningless. Conversely, an intervention that reduces cardiac events by 30% might not reach statistical significance in a small pilot study — but the effect size tells you it is worth studying with adequate power.
Large Samples and Tiny P-Values
As sample size n grows, the standard error shrinks as 1/√n, making test statistics larger and p-values smaller. With enough data, any deviation from the null hypothesis — no matter how trivial — becomes detectable. This means that in large datasets, significant p-values are nearly guaranteed, even for effects of no practical consequence.
The test statistic for a one-sample z-test illustrates the problem directly:
The antidote is to always report effect sizes alongside p-values, and to pre-specify the minimum effect size that would be practically meaningful before collecting data. This prevents post-hoc rationalization of trivial effects as discoveries.
Every hypothesis test result should report three things: the p-value (evidence against H₀), the effect size (magnitude of the effect), and a confidence interval (uncertainty range). A p-value alone tells only a fraction of the story. The American Statistical Association’s 2016 statement on p-values made exactly this recommendation.
The Reproducibility Crisis
The cumulative effect of p-hacking, selective publication, and inadequate sample sizes came into sharp focus in the early 2010s when large-scale replication efforts found that many published findings could not be reproduced. The Open Science Collaboration’s 2015 paper replicated 100 psychology studies and found that only 36% showed statistically significant results in the replication, compared to 97% in the original papers.
This reproducibility crisis is not confined to psychology. Similar replication failures have appeared in social science, nutrition research, cancer biology, preclinical pharmacology, and economics. The root causes are well understood: publication bias (journals favor positive results), small sample sizes that give noisy effect estimates, researcher degrees of freedom without correction, and insufficient transparency in methods reporting.
The landmark “False Positive Psychology” paper showed that with just four researcher degrees of freedom — adding covariates, collecting more data after peeking, choosing among outcome variables, or dropping conditions — the probability of a false positive could reach 61% even while honestly following conventional rules. Flexibility without pre-specification is the enemy of valid inference.
Pre-Registration: The Solution
Pre-registration is the practice of publicly committing to your hypotheses, sample size, outcome variables, and analysis plan before collecting data. By locking these choices in advance, researchers eliminate the degrees of freedom that enable p-hacking and make the distinction between confirmatory and exploratory analysis explicit.
Pre-registration does not prohibit exploration — it just requires you to label exploratory analyses as such. A pre-registered study that reports p = 0.04 carries much stronger evidential weight than an unregistered study reporting the same value, because the former cannot have been shaped by knowledge of the data.
Platforms like the Open Science Framework (OSF), AsPredicted.org, and the American Economic Association’s RCT registry make pre-registration easy and free. Registered Reports — where journals peer-review and accept studies based on their design before results are known — extend pre-registration’s benefits by eliminating publication bias at its source.
1. Pre-register hypotheses and analysis plan before data collection. 2. Report effect sizes and CIs alongside every p-value. 3. Apply multiple comparisons corrections when testing more than one hypothesis. 4. Distinguish confirmatory from exploratory analyses in the write-up. 5. Share data and code to enable independent verification. 6. Power your study adequately (≥ 80% power at the smallest effect that matters) before starting.
- P-hacking inflates Type I error by exploiting researcher degrees of freedom: stopping early, testing multiple outcomes, or choosing analyses after seeing the data.
- Multiple comparisons: running m tests at α each gives FWER = 1 − (1−α)ᵐ. With 20 tests at 5%, the chance of a false positive exceeds 60%.
- Bonferroni correction controls FWER by testing at α/m. Conservative but safe for small m or high-stakes tests.
- Benjamini-Hochberg controls the False Discovery Rate. More powerful than Bonferroni when many true effects exist (genomics, neuroimaging).
- Statistical ≠ practical significance. A tiny effect in a huge sample can yield p < 0.001. Always report Cohen’s d or equivalent effect size.
- Large samples guarantee small p-values even for meaningless effects. Effect size is sample-size-independent; p-values are not.
- Reproducibility crisis: decades of p-hacking and publication bias have filled the literature with false positives. Many landmark findings do not replicate.
- Pre-registration eliminates researcher degrees of freedom by committing hypotheses and analysis plan before data collection. It is the gold standard for confirmatory research.