The Audacious Idea
Classical inference — z-intervals, t-intervals — works by deriving the exact or approximate sampling distribution of a statistic under specific model assumptions. But what do you do when the statistic is a median, a ratio of variances, a correlation coefficient, or something even more exotic? Deriving a sampling distribution by hand quickly becomes intractable or requires assumptions you cannot verify.
In 1979, statistician Bradley Efron proposed a remarkably simple idea: if you do not know the population distribution, use the sample itself as a stand-in. Resample from your data with replacement, compute the statistic on each resample, and let the resulting distribution of resampled statistics approximate the true sampling distribution. He called this the bootstrap — named after the phrase “to pull oneself up by one’s bootstraps,” conjuring the image of lifting yourself with your own data.
This lesson covers the bootstrap algorithm, how it produces confidence intervals, why it works for virtually any statistic, the distinction between parametric and nonparametric bootstrap, and where the method has known limitations.
The Bootstrap Algorithm
Suppose we have a sample x = (x1, x2, …, xn) from some unknown distribution F. We want a confidence interval for a parameter θ estimated by statistic T(x) (e.g., the sample mean, median, or correlation).
Step 1: Draw a bootstrap sample x* = (x*1, …, x*n) by sampling n observations from x with replacement.
Step 2: Compute the bootstrap statistic T* = T(x*).
Step 3: Repeat steps 1–2 a large number of times B (typically B = 1,000 to 10,000).
Step 4: Use the B values T*1, T*2, …, T*B to estimate the sampling distribution of T.
Because each bootstrap sample is drawn with replacement from n observations, some original observations appear multiple times and others not at all. On average, about 63.2% of distinct observations appear in any given bootstrap sample (the rest are duplicates of already-selected observations).
Bootstrap Standard Error
The most basic use of bootstrap resamples is to estimate the standard error of any statistic — the quantity that typically requires analytical derivation in classical statistics.
This is liberating: for any statistic — no matter how unusual — the standard error is just the standard deviation of its bootstrap distribution. You never need to derive a formula.
Bootstrap Confidence Intervals
There are several ways to convert bootstrap resamples into confidence intervals. The two most common are the percentile method and the basic (pivot) method.
Percentile interval: Simply take the α/2 and 1 − α/2 quantiles of the bootstrap distribution as the confidence interval boundaries:
Basic interval (reflection method): Instead of using the percentiles of T* directly, use them to correct for bias and create a reflected interval:
More advanced variants include the BCa interval (bias-corrected and accelerated), which adjusts for both bias and skewness in the bootstrap distribution and is the method most recommended in modern practice. For large B and roughly symmetric bootstrap distributions, all methods give similar results.
Bootstrap for Any Statistic
The bootstrap’s most powerful feature is its universality. Classical parametric methods provide formulas for the mean and proportion, but for virtually everything else, the bootstrap is the go-to approach. Here are key examples:
Median: No closed-form SE exists for non-normal distributions — bootstrap provides one effortlessly.
Correlation coefficient: The Fisher z-transform gives a classical CI, but bootstrap handles non-normality and small samples better.
Ratio of means: No simple formula; bootstrap works directly.
Regression coefficients: Bootstrap CIs are more robust than classical CIs when residuals are non-normal or heteroscedastic.
Machine learning model metrics: AUC, accuracy, F1-score — bootstrap is the standard way to attach uncertainty to these.
Parametric vs. Nonparametric Bootstrap
The algorithm described above is the nonparametric bootstrap — it resamples from the empirical distribution of the data, making no assumptions about the underlying population shape. This is the most commonly used variant.
The parametric bootstrap instead fits a parametric model to the data (e.g., assumes the data are normal with estimated mean and variance), then generates bootstrap samples by drawing from the fitted model rather than from the original data. This is useful when:
• You have strong prior knowledge of the distributional family.
• Your sample is small and the empirical distribution is a poor approximation of F.
• You need bootstrap samples that extend beyond the observed data range (the nonparametric bootstrap can never produce values outside the original sample).
Nonparametric: resample from the data. No distributional assumption. Universally applicable. Requires n to be large enough that the empirical distribution well-represents F.
Parametric: resample from a fitted model. Leverages distributional knowledge for greater efficiency. Fails badly when the assumed model is wrong.
Why the Bootstrap Works
The theoretical justification hinges on the plug-in principle: replace the unknown population distribution F with the empirical distribution F̂ (which assigns probability 1/n to each observed data point). Any functional of F can then be estimated by the same functional of F̂.
As n → ∞, F̂ converges to F by the Glivenko–Cantelli theorem. The bootstrap distribution of T* − T converges to the true sampling distribution of T − θ. So bootstrap confidence intervals are asymptotically valid — they achieve the correct coverage probability as n grows.
For finite samples, the quality of bootstrap approximation depends on how well F̂ represents F. This is why small samples remain a challenge: with n = 10 there are only about 10 distinct observations in the empirical distribution, which may be a poor stand-in for the true continuous distribution.
When Bootstrap Fails
The bootstrap is not a magic wand. There are known failure modes:
Extreme quantiles and max/min: Estimating the maximum of a distribution fails — the bootstrap maximum can never exceed the observed maximum.
Heavy-tailed distributions: When variance is infinite (e.g., Cauchy), the bootstrap may not converge correctly.
Very small samples (n < 15): The empirical distribution is a crude approximation of F; coverage can be unreliable.
Non-i.i.d. data: Time series, spatial data, and clustered data require special bootstrap variants (block bootstrap, cluster bootstrap) that preserve the dependence structure.
Model selection and sparsity: Bootstrapping a model selection procedure (e.g., LASSO) can give misleading results because the discrete nature of selection is not well-captured by resampling.
Computational Implementation
The bootstrap’s main cost is computational: you must run your statistic calculation B times. With B = 1,000 and a fast statistic (e.g., median of n = 100 observations), this is milliseconds on a modern computer. With B = 10,000 and a computationally intensive statistic (e.g., a cross-validated AUC on a neural network), it can take hours.
In Python, a bootstrap CI for any statistic is straightforward:
import numpy as np
rng = np.random.default_rng(42)
B = 5000
boot_stats = [np.median(rng.choice(x, n, replace=True)) for _ in range(B)]
ci = np.percentile(boot_stats, [2.5, 97.5])
SciPy 1.7+ includes scipy.stats.bootstrap which implements percentile, basic, and BCa methods. R’s boot package is the reference implementation, supporting all major variants and accelerated computation.
- Core idea: treat the sample as a stand-in for the population; resample with replacement B times to approximate the sampling distribution of any statistic.
- Bootstrap SE: standard deviation of the B bootstrap statistics — requires no analytical formula.
- Percentile CI: [T*(α/2), T*(1−α/2)] — the simplest and most widely used bootstrap CI.
- Universality: works for median, correlation, regression coefficients, AUC, and virtually any other estimable quantity.
- Nonparametric bootstrap: resamples from empirical distribution; no model assumption needed; valid asymptotically.
- Parametric bootstrap: resamples from fitted model; more efficient when the model is correct; misleading when the model is wrong.
- Known failures: extreme-order statistics, heavy tails, very small n, non-i.i.d. data (use block/cluster bootstrap instead).