Where Statistics Goes Next
The modules of this course have equipped you with a rigorous foundation: data types and sampling, data visualization, probability, distributions, estimation, hypothesis testing, regression, time series, and experimental design. These are the core tools that working statisticians and data scientists use daily. But the field extends far beyond them.
This lesson surveys six frontiers — survival analysis, spatial statistics, functional data analysis, high-dimensional statistics, causal inference, and the intersection with machine learning. Each is a mature sub-discipline with its own literature, software ecosystem, and community. The goal here is to give you enough of a map that you can decide which frontier to explore next, and know what to read when you get there.
Survival Analysis
Survival analysis deals with time-to-event data: how long until a patient dies, a machine fails, a customer churns, or a user clicks. The defining challenge is censoring: many subjects have not experienced the event by the end of the study. A patient who is still alive at the study’s conclusion contributes information — we know they survived at least this long — but their survival time is unknown. Discarding censored observations would introduce severe bias; survival analysis handles them correctly.
The central object is the survival function S(t) = P(T > t): the probability that the event has not yet occurred by time t. Its non-parametric estimator is the Kaplan–Meier estimator, which produces a step function that drops at each observed event time.
The Cox proportional hazards model extends survival analysis to covariates. It models the hazard — the instantaneous rate of the event given survival to time t — as a baseline hazard multiplied by an exponential function of covariates. Its key insight is that regression coefficients can be estimated without specifying the baseline hazard, making it semi-parametric. Cox regression is one of the most widely cited statistical methods in the biomedical literature.
Censoring — event not observed by end of study; the observation is incomplete but informative.
Hazard function h(t) — instantaneous event rate given survival to time t.
Log-rank test — non-parametric comparison of survival curves between groups.
Software: R packages survival and survminer; Python’s lifelines and scikit-survival.
Spatial Statistics
Spatial statistics analyzes data that have a geographic component. Standard statistical methods assume independence between observations. Spatial data violates this: nearby locations tend to be more similar than distant ones — a phenomenon called spatial autocorrelation. Ignoring it leads to underestimated standard errors and overconfident inferences.
Tobler’s First Law of Geography states: “everything is related to everything else, but near things are more related than distant things.” Quantifying this is done via the variogram, which plots the average squared difference between values at pairs of locations as a function of their distance (the lag). The variogram rises from a nugget at zero lag (measurement error and micro-scale variation) to a sill (the total variance) over a range (the distance at which spatial correlation vanishes).
Kriging, developed by the mining engineer Danie Krige, is a geostatistical technique for optimal spatial prediction. Given observed values at some locations, kriging produces a best linear unbiased predictor at any unobserved location, along with a prediction variance. It is the spatial analogue of Gaussian process regression.
Beyond geostatistics, spatial statistics includes areal data analysis (data aggregated over administrative regions like counties) using tools like the Moran’s I statistic and conditional autoregressive (CAR) models, and point process analysis (modeling the locations of events like disease cases or earthquake epicenters) using tools from the theory of spatial point processes.
Functional Data Analysis
Functional data analysis (FDA) treats curves, surfaces, or other functions as the unit of observation, rather than a vector of scalar measurements. A growth curve for an individual child, a temperature profile recorded over a year, or an EEG trace for a single trial — each is a function, and FDA provides tools to summarize, compare, and model them.
The mean function and covariance function replace the mean vector and covariance matrix. Functional principal component analysis (FPCA) decomposes curves into orthonormal basis functions (eigenfunctions of the covariance operator) and scalar scores, reducing infinite-dimensional data to a manageable set of features. This is the functional analogue of PCA.
FDA is especially useful in biomedical research (growth curves, longitudinal biomarkers), environmental science (temperature and precipitation time series), and engineering (sensor signals over time). The key insight is that treating the entire curve as the unit of analysis captures information that scalar summaries (e.g., the maximum or the area under the curve) discard.
Classical statistics: observe a vector x ∈ ℝp.
Functional statistics: observe a function x(t) defined on an interval [a, b]. The “dimension” is infinite, but structure (smoothness) makes the problem tractable.
FPCA: x(t) ≈ μ(t) + Σk ξk φk(t), where ξk are random scoresHigh-Dimensional Statistics
High-dimensional statistics addresses problems where the number of variables p is comparable to or exceeds the number of observations n. Classical asymptotic theory assumed n → ∞ with p fixed; the modern regime is n, p → ∞ together, often with p ≫ n. Genomics is the canonical example: a single study might have n = 200 patients and p = 20,000 gene expression features.
Ordinary least squares collapses when p ≥ n: the design matrix is rank-deficient, coefficients are not identified, and out-of-sample prediction is catastrophic. Regularization rescues the situation by adding a penalty term to the least-squares objective that shrinks or eliminates coefficients.
Ridge regression uses an L2 penalty that shrinks all coefficients toward zero but rarely sets them exactly to zero. Elastic net combines L1 and L2 penalties to achieve both variable selection and handling of correlated predictors. The LASSO (Least Absolute Shrinkage and Selection Operator) has become particularly prominent because its sparsity-inducing property makes the result interpretable.
Beyond regression, high-dimensional statistics encompasses multiple testing (controlling the false discovery rate across thousands of simultaneous tests), covariance estimation (sparse or structured covariance matrices when p ≫ n), and compressed sensing (recovering sparse signals from far fewer measurements than the signal dimension).
Causal Inference Deep Dive
Module 11 introduced causal inference in the context of experimental design. Here we go deeper into the potential outcomes framework (also called the Rubin Causal Model), which provides the formal language for causal questions. For each unit i, define Yi(1) as the outcome if treated and Yi(0) as the outcome if untreated. The individual treatment effect is Yi(1) − Yi(0), but this is fundamentally unobservable because each unit can only be in one condition at a time — the fundamental problem of causal inference.
The average treatment effect (ATE) = E[Y(1) − Y(0)] is estimable from randomized experiments, where treatment assignment is independent of potential outcomes. In observational studies, this independence fails; we must control for confounders through matching, propensity score weighting, or regression adjustment.
Instrumental variables (IV) offer an elegant solution when an unobserved confounder exists: find a variable Z that affects treatment assignment but affects the outcome only through treatment. Z is the instrument. IV estimation — particularly two-stage least squares (2SLS) — identifies a causal effect for the complier subpopulation (those whose treatment status changes when Z changes).
Regression discontinuity (RD) exploits sharp cutoffs in assignment rules. If students scoring above 70 receive a scholarship and those below do not, we can compare outcomes for students just above and just below the threshold — groups that are nearly identical except for the scholarship — to estimate its causal effect. Difference-in-differences (DiD) compares changes over time between a treatment group and a control group, identifying causal effects under a parallel trends assumption.
Randomized experiment — Gold standard; treatment independent of potential outcomes by design.
Matching / IPW — Observational; controls confounders assumed to be measured.
Instrumental variables — Needs an exogenous instrument; identifies LATE for compliers.
Regression discontinuity — Needs a sharp assignment cutoff; local causal effect near the threshold.
Difference-in-differences — Needs parallel trends; leverages panel data.
Machine Learning and Statistics: The Intersection
Statistics and machine learning grew from different traditions — statistics from mathematics and experimental science, machine learning from computer science and artificial intelligence — but they have been converging rapidly. Understanding the relationship illuminates both fields.
Prediction vs. inference is the clearest dividing line. Classical statistics emphasizes inference: estimating parameters, quantifying uncertainty, testing hypotheses. Machine learning emphasizes prediction: building models that generalize to new data, evaluated by held-out test error. A linear regression model is transparent and inferentially rich; a gradient boosted forest may predict better but is harder to interpret. Neither goal is superior; they serve different purposes.
Regularization is the bridge. Ridge regression = L2-penalized linear regression = a special case of Gaussian process regression with a specific prior. LASSO = L1-penalized linear regression = MAP estimation under a Laplace prior. Cross-validation for lambda selection = empirical Bayes marginal likelihood estimation. The same procedures arise independently in both traditions.
Conformal prediction is a modern framework that wraps any black-box predictor — a neural network, a random forest, any algorithm — in a valid prediction interval with finite-sample coverage guarantees, without parametric assumptions. It extends statistical uncertainty quantification to machine learning models that classical theory does not cover.
Double machine learning (DML), developed by Chernozhukov and colleagues, combines high-dimensional machine learning with causal inference. The idea: use ML to flexibly partial out the effects of high-dimensional controls on both the treatment and the outcome, then estimate causal effects on the residuals. This achieves near-parametric rates of convergence for causal parameters even when the nuisance functions (the relationship between controls and outcome) are high-dimensional.
The synthesis is still unfolding. Modern statisticians need ML fluency; modern ML practitioners benefit from statistical discipline around uncertainty, causal reasoning, and valid inference. The boundaries are dissolving, and the most interesting work sits exactly at the intersection.
- Survival analysis handles time-to-event data with censoring. The Kaplan–Meier estimator describes survival curves non-parametrically; Cox regression extends this to covariates semi-parametrically.
- Spatial statistics accounts for geographic dependence. Kriging provides optimal spatial prediction; variograms quantify how correlation decays with distance.
- Functional data analysis treats entire curves as observations. FPCA decomposes curves into interpretable components, enabling regression and clustering of functional data.
- High-dimensional statistics addresses p ≫ n settings. LASSO regularization simultaneously selects variables and estimates coefficients; false discovery rate control handles massive multiple testing.
- Causal inference extends beyond observational correlation using potential outcomes, instrumental variables, regression discontinuity, and difference-in-differences.
- Statistics and machine learning are converging. Regularization, cross-validation, and uncertainty quantification are shared foundations; conformal prediction and double machine learning exemplify productive synthesis.