The Problem of Estimation
At the heart of statistical inference lies a fundamental challenge: we want to know something about a population — its mean, its variance, the probability of some event — but we can only observe a finite sample. Point estimation is the art of using sample data to produce a single best guess for an unknown population parameter.
The unknown parameter might be the mean μ of a normal distribution, the success probability p of a Bernoulli trial, the rate λ of a Poisson process, or any other quantity that characterizes how data are generated. We cannot observe these parameters directly; we can only observe their noisy consequences in the data.
In this lesson we build the vocabulary of estimation (estimator vs. estimate, bias, consistency, efficiency), develop two systematic estimation frameworks (method of moments and maximum likelihood), and explore the fundamental bias-variance trade-off that governs all estimation problems.
Estimator vs. Estimate
Before doing anything else, we need to be precise about language. An estimator is a function of the data — a rule or procedure that maps a sample to a number. An estimate is the specific value that an estimator produces when applied to a specific dataset.
For example, the sample mean X̄ = (X1 + X2 + … + Xn) / n is an estimator of the population mean μ. If we observe the specific values {2.1, 3.4, 1.8, 4.0}, then x̄ = 2.825 is the estimate — the particular number produced by this rule on this dataset.
The estimator is a random variable: it has a distribution (the sampling distribution) because it depends on random data. The estimate is a fixed number, the realization of that random variable. This distinction matters enormously: when we talk about properties like bias and variance, we are talking about the estimator — its behavior across all possible samples — not about any single estimate.
Estimator: Θ̂ = g(X1, …, Xn) — a function of random variables, itself a random variable with a distribution.
Estimate: θ̂ = g(x1, …, xn) — the specific number computed from observed data. Lower case = observed; upper case = random variable.
Unbiasedness
The most intuitive desirable property for an estimator is unbiasedness: on average, across all possible samples, the estimator should equal the true parameter value. Formally:
The sample mean X̄ is an unbiased estimator of μ for any population with a finite mean. The sample variance S² = ∑(Xi − X̄)² / (n − 1) is an unbiased estimator of σ² — the divisor n − 1 rather than n is precisely what makes it unbiased.
Biasedness, however, is not always fatal. Some estimators are deliberately biased to reduce variance, yielding better overall performance. The sample variance using divisor n (the MLE of σ²) is biased, but only slightly so, and it is still useful in many contexts. The right criterion depends on how you weigh bias against variance.
Consistency and Efficiency
Consistency is a large-sample property: an estimator is consistent if it converges in probability to the true parameter as n → ∞. Formally, Θ̂n is consistent for θ if for any ε > 0:
The Law of Large Numbers guarantees that the sample mean is consistent for any population with a finite mean. More generally, any unbiased estimator whose variance goes to zero as n increases is consistent.
Efficiency compares estimators within the class of unbiased ones. An unbiased estimator is efficient if it achieves the smallest possible variance among all unbiased estimators. The Cramér-Rao Lower Bound (CRLB) provides the theoretical minimum variance any unbiased estimator can attain:
Method of Moments
The method of moments (MOM) is one of the oldest and most intuitive estimation techniques. The idea is simple: the k-th population moment is E[Xk], which depends on the unknown parameter(s) θ. The k-th sample moment is mk = (1/n)∑Xik, which can be computed from data. The MOM estimator sets sample moments equal to population moments and solves for the parameters.
For a distribution with a single parameter, we use the first moment. For two parameters, we use the first and second moments, and so on. The system of equations is:
Example — Normal distribution: The Normal N(μ, σ²) has E[X] = μ and E[X²] = σ² + μ². Setting m1 = X̄ gives μ̂ = X̄, and solving from the second moment gives σ̂² = (1/n)∑(Xi − X̄)². Simple and intuitive.
The method of moments is easy to apply and the resulting estimators are consistent. However, they may not be the most efficient — they do not use all the information in the data in the most effective way. For that, we turn to maximum likelihood.
Maximum Likelihood Estimation
Maximum Likelihood Estimation (MLE) is the dominant estimation paradigm in modern statistics. The idea is to find the parameter value that makes the observed data most probable — the parameter under which the data are “most likely.”
Given observations x1, …, xn from a density (or pmf) f(x; θ), the likelihood function treats the data as fixed and views f as a function of θ:
Because products of many small numbers are numerically problematic, we almost always work with the log-likelihood ℓ(θ) = log L(θ). Since log is monotone increasing, maximizing ℓ and maximizing L give the same answer:
To find the MLE, differentiate the log-likelihood with respect to θ and set the derivative to zero. This gives the score equation:
MLE for Common Distributions
Let’s work through the three most important examples, which illustrate the MLE technique and produce results you will use constantly.
Normal distribution N(μ, σ²): The log-likelihood of n i.i.d. normal observations is:
Consistency: Under mild regularity conditions, the MLE θ̂ converges in probability to the true θ as n → ∞.
Asymptotic efficiency: The MLE achieves the Cramér-Rao bound asymptotically — no consistent estimator has smaller asymptotic variance.
Invariance: If θ̂ is the MLE of θ, then g(θ̂) is the MLE of g(θ) for any function g. This makes transformations trivial.
The Bias-Variance Trade-off
Unbiasedness and small variance are both desirable, but they often conflict. The mean squared error (MSE) combines them into a single criterion that measures the overall accuracy of an estimator:
Consider estimating the normal variance σ². The MLE σ̂² = (1/n)∑(Xi − X̄)² is biased (E[σ̂²] = ((n-1)/n)σ²), but it has smaller variance than the unbiased S² = (1/(n-1))∑(Xi − X̄)². For small n, this bias-variance exchange can be worthwhile; for large n, both estimators are nearly equivalent.
The bias-variance trade-off is not just a curiosity — it is the central tension in all of statistics and machine learning. Regularization (Ridge, Lasso) deliberately introduces bias to reduce variance. Shrinkage estimators (like the James-Stein estimator) intentionally shrink toward zero or some prior value, trading unbiasedness for lower MSE in high-dimensional settings.
Unbiased, high variance: Darts centered on the bullseye but scattered widely. On average correct, but no single dart is close.
Biased, low variance: Darts clustered tightly but offset from bullseye. Consistently wrong in the same direction, but precise.
Unbiased, low variance: The ideal — centered on bullseye and tightly grouped. Achievable only if there is an efficient unbiased estimator (UMVUE).
- Estimator vs. estimate: an estimator Θ̂ = g(X1,…,Xn) is a random variable (a rule); an estimate θ̂ = g(x1,…,xn) is the specific number it produces on observed data.
- Unbiasedness: E[Θ̂] = θ for all θ. The sample mean is unbiased for μ; the sample variance with divisor n−1 is unbiased for σ².
- Consistency: Θ̂n → θ in probability as n → ∞. Any unbiased estimator with variance → 0 is consistent.
- Efficiency: the Cramér-Rao lower bound gives the minimum variance achievable by any unbiased estimator. The MLE is asymptotically efficient.
- Method of moments: set sample moments equal to population moments and solve. Simple, consistent, but not always most efficient.
- MLE: choose θ to maximize L(θ) = ∏f(xi;θ), equivalently maximize ℓ(θ) = ∑log f(xi;θ). MLEs are consistent, asymptotically efficient, and invariant to reparameterization.
- Bias-variance trade-off: MSE = Var + Bias². Introducing bias can reduce MSE if it sufficiently shrinks variance. This tension is fundamental to all of statistics and machine learning.