The Hardest Question in Statistics
Statistics is full of hard questions, but one towers above the rest: does X cause Y? Correlation is easy to measure. Causation is notoriously difficult to establish. Ice cream sales correlate with drowning deaths — not because ice cream causes drowning, but because a third variable, hot weather, drives both. This is the essence of the causal inference problem: disentangling true cause-and-effect relationships from the web of correlations that real-world data inevitably produces.
The previous two lessons covered tools for well-designed experiments. But the world does not always cooperate with our experimental plans. We cannot randomly assign people to smoking or non-smoking. We cannot randomly assign countries to different economic policies. We cannot always wait for a randomized trial before making a decision. Causal inference gives us principled tools for reasoning about cause and effect in precisely these situations — when randomization is impossible, unethical, or too slow.
Correlation measures the statistical association between two variables. It is symmetric: X correlates with Y exactly as much as Y correlates with X.
Causation is asymmetric and directional: X causes Y means that intervening to change X would change Y, but not vice versa.
Correlation ≠ Causation — but with the right design, we can get close.Confounding Variables
A confounder is a variable that influences both the treatment and the outcome, creating a spurious association between them. In the ice cream and drowning example, hot weather is the confounder: it causes more ice cream consumption and more swimming, which leads to more drowning incidents. If we naively estimated the effect of ice cream on drowning, we would find a positive association that is entirely driven by the confounder.
Confounders are not just theoretical curiosities. They are the central obstacle to causal inference from observational data. Studies have found that people who carry lighters are more likely to develop lung cancer — not because lighters cause cancer, but because lighter carriers are more likely to smoke. Hormone replacement therapy appeared protective against heart disease in early observational studies, but randomized trials later showed it could increase risk — the association was driven by confounding, because healthier, wealthier women were more likely to use HRT and also less likely to develop heart disease.
Let T = treatment, Y = outcome, C = confounder. We observe P(Y | T), but we want P(Y | do(T)) — the probability of Y when we intervene to set T, rather than merely observe T. If C affects both T and Y, then P(Y | T = 1) ≠ P(Y | do(T = 1)).
Simpson’s Paradox
Simpson’s paradox is one of the most counterintuitive results in statistics: a trend that appears in several groups of data reverses when the groups are combined. It is not a logical contradiction — it is a vivid demonstration of how confounding can completely reverse apparent relationships.
The classic example: a hospital compares two doctors. Doctor A has an 80% survival rate, Doctor B has a 90% survival rate. You should go to Doctor B, right? But suppose Doctor A specializes in complex cases (90% survival for complex, 50% for simple) and Doctor B handles routine cases (95% for complex, 70% for simple). When you look at each case type separately, Doctor A outperforms Doctor B. But Doctor A’s overall rate looks worse because they see more complex patients. The confounder — case complexity — is driving the reversal.
Simpson’s paradox matters because it shows that aggregate statistics can be genuinely misleading. The reversal is not a statistical artifact — it reflects real structure in the data that you miss when you ignore the confounder. The correct analysis stratifies by the confounder, computing effects within each stratum rather than collapsing across them.
Randomized Controlled Trials: The Gold Standard
The reason the randomized controlled trial (RCT) holds the title of gold standard for causal inference is simple: randomization eliminates confounding by design. When we randomly assign individuals to treatment or control, the groups are expected to be identical on average across all variables — both measured and unmeasured. Any difference in outcomes can therefore be attributed to the treatment alone.
This is the average treatment effect (ATE): the expected difference in outcomes between treatment and control conditions, averaged across the population.
The elegance of randomization lies in the potential outcomes framework developed by Donald Rubin. Each unit has two potential outcomes: what would happen under treatment, Y(1), and what would happen under control, Y(0). We can only ever observe one of these — this is called the fundamental problem of causal inference. Randomization lets us estimate the average treatment effect by comparing groups, even though we can never observe both potential outcomes for the same individual.
Observational Studies: Matching and Propensity Scores
When randomization is not possible, we must try to adjust for confounders using the data we have. The core idea is to compare treated and control units that look similar on relevant background variables — essentially creating an artificial experiment from observational data.
Exact matching pairs each treated unit with a control unit that has identical values on all confounders. This works well when there are few confounders and large samples, but quickly becomes infeasible as the number of confounders grows — the curse of dimensionality means that exact matches become increasingly rare.
The propensity score, introduced by Rosenbaum and Rubin in 1983, solves this problem elegantly. The propensity score is the probability that a unit receives treatment given its observed covariates:
The key theorem: if we have measured all the confounders (the ignorability assumption), then conditioning on the propensity score is sufficient to remove confounding. In practice, the propensity score is estimated from data using logistic regression or a machine learning model, then used in one of three ways: propensity score matching (pair units with similar scores), stratification (group units into bins by score and estimate effects within each bin), or inverse probability weighting (weight each unit by the inverse of its propensity score).
Difference-in-Differences
Difference-in-differences (DiD) is a powerful quasi-experimental design that exploits the timing of treatment to identify causal effects. The idea is to compare the change in outcomes over time for a treated group to the change in outcomes over the same period for a control group. If the control group serves as a valid counterfactual for what would have happened to the treated group in the absence of treatment, then the difference between these two changes isolates the causal effect.
A classic application: Card and Krueger’s 1994 study of the minimum wage. New Jersey raised its minimum wage while neighboring Pennsylvania did not. If minimum wage increases truly reduce employment (as classical theory predicts), employment in fast food restaurants should fall in New Jersey relative to Pennsylvania. The DiD estimate showed no such decline — a result that changed the economics debate on minimum wages.
The critical assumption for DiD validity is the parallel trends assumption: in the absence of treatment, the treated and control groups would have followed the same trend. This cannot be tested directly (we never observe the counterfactual trend), but pre-treatment trends can provide supporting evidence. If the two groups were trending parallel before treatment, this makes the parallel trends assumption more plausible.
Instrumental Variables
Even propensity score methods and DiD cannot handle unmeasured confounders. If there is an unobserved variable that affects both treatment and outcome, these methods will still produce biased estimates. The instrumental variables (IV) approach offers a way out in certain situations.
An instrument is a variable Z that satisfies three conditions: (1) Z is correlated with the treatment T (relevance), (2) Z affects the outcome Y only through its effect on T, not directly (exclusion restriction), and (3) Z is independent of unmeasured confounders (exogeneity). The instrument shifts treatment assignment in a way that is “as good as random,” allowing us to use the variation in treatment caused by Z to estimate the causal effect of T on Y.
Draft lottery number as an instrument for military service (to estimate the effect of service on earnings). The lottery number is random, affects service probability, and affects earnings only through service.
Distance to college as an instrument for education (to estimate the effect of education on wages). Distance affects whether you go to college but does not directly affect your earnings except through education.
Rainfall as an instrument for economic shocks (in development economics). Rainfall affects income in agricultural economies but does not directly affect the political or social outcome being studied.
The IV estimator divides the reduced-form effect (the effect of Z on Y) by the first-stage effect (the effect of Z on T). Intuitively, we scale the impact of the instrument on the outcome by how much it shifts treatment, to recover the treatment effect. IV estimates apply specifically to compliers — units whose treatment status is shifted by the instrument — so the IV estimate is technically a local average treatment effect (LATE), not the population ATE.
The Hierarchy of Evidence
Not all evidence is equally credible for causal claims. Epidemiologists and economists have developed informal hierarchies ranking the strength of different study designs for causal inference.
At the top sits the randomized controlled trial: randomization eliminates confounding by design, making causal claims robust. Just below are quasi-experiments that exploit natural or policy-driven variation — difference-in-differences, regression discontinuity, and instrumental variables — which can approximate experimental variation without randomizing.
Lower in the hierarchy are propensity score methods and matching, which require the ignorability assumption (no unmeasured confounders). At the bottom sit unadjusted observational associations, which are easily confounded and should not be interpreted causally without strong additional justification.
The key lesson is not that observational data is useless for causal inference — it is indispensable when experiments are infeasible. But the credibility of a causal claim depends on the quality of the design and the plausibility of the required assumptions. Transparency about those assumptions, and sensitivity analyses that test them, are the hallmarks of careful causal reasoning.
1. State the causal question precisely — what is the treatment, what is the outcome, and what population?
2. Identify the confounders — draw a causal diagram (DAG) to reason about what must be adjusted for.
3. Choose the right design — RCT if possible; quasi-experiment if natural variation exists; propensity methods otherwise.
4. State assumptions explicitly — ignorability? Parallel trends? Instrument validity? These cannot be verified from data alone.
5. Test robustness — sensitivity analyses show how much hidden bias would need to exist to overturn your conclusion.
- Correlation measures association; causation requires that an intervention on X would change Y. Confounders create spurious associations by influencing both treatment and outcome.
- Simpson’s paradox shows how an aggregate trend can reverse within subgroups when a confounder is ignored — always stratify by relevant variables before drawing causal conclusions.
- Randomized controlled trials eliminate confounding by design. The potential outcomes framework formalizes why: randomization makes treated and control groups comparable in expectation on all covariates, measured and unmeasured.
- Propensity score methods — matching, stratification, and inverse probability weighting — can adjust for observed confounders in observational studies. They require the ignorability assumption: no unmeasured confounders after conditioning on the observed covariates.
- Difference-in-differences compares trends over time between treated and control groups, removing time-invariant confounding. Its validity rests on the parallel trends assumption.
- Instrumental variables use an exogenous source of variation in treatment to estimate causal effects even in the presence of unmeasured confounders. The instrument must be relevant, satisfy the exclusion restriction, and be independent of unmeasured confounders.