STAT 101
M05 · L01
Module 5 · Lesson 1
Point Estimation
We observe data. We want to know the underlying parameter. Point estimation gives us a single best guess — and the tools to know how good that guess is.
01 / 13
STAT 101
M05 · L01
Precision of language
Estimator vs. Estimate
Estimator
g(X₁,…,Xₙ)
random variable
random variable
Estimate
g(x₁,…,xₙ)
specific number
specific number
Upper case = random variable. Lower case = observed value. Properties like bias and variance belong to the estimator, not any single estimate.
02 / 13
STAT 101
M05 · L01
Desirable property #1
Unbiasedness
On average, over all possible samples, the estimator hits the true target:
Unbiasedness
E[\hat{\theta}] = \theta
Example
X̄ is unbiased for μ.
S² with n−1 is unbiased for σ².
S² with n−1 is unbiased for σ².
03 / 13
STAT 101
M05 · L01
Desirable property #2
Consistency
More data = closer to the truth. Formally, θ̂ₙ converges in probability to θ:
Consistency
\lim_{n\to\infty} P(|\hat{\theta}_n - \theta| > \varepsilon) = 0
04 / 13
STAT 101
M05 · L01
Desirable property #3
Efficiency & CRLB
- Among all unbiased estimators, prefer the lowest variance
- The Cramér-Rao bound sets the minimum achievable variance
- An estimator hitting the bound is called UMVUE
- The MLE is asymptotically efficient
05 / 13
STAT 101
M05 · L01
Estimation framework #1
Method of Moments
Set sample moments equal to population moments, then solve:
MOM
E[X^k] = \tfrac{1}{n}\sum_{i=1}^n X_i^k
Normal example
μ̂ = X̄, σ̂² = (1/n)Σ(Xᵢ−X̄)²
06 / 13
STAT 101
M05 · L01
Estimation framework #2
Maximum Likelihood
Choose θ to make the observed data as probable as possible:
Likelihood
L(\theta) = \prod_{i=1}^{n} f(x_i;\,\theta)
In practice, maximize the log-likelihood ℓ(θ) = Σ log f(xᵢ; θ).
07 / 13
STAT 101
M05 · L01
Solving for the MLE
The Score Equation
Differentiate log-likelihood, set to zero:
Score equation
\dfrac{\partial\,\ell(\theta)}{\partial\theta} = 0 \implies \hat{\theta}_{\mathrm{MLE}}
Why log?
Products → sums. Numerically stable. Same maximizer.
08 / 13
STAT 101
M05 · L01
MLE in practice
MLE for the Normal
- Maximize over μ → μ̂ = X̄ (sample mean)
- Maximize over σ² → σ̂² = (1/n)Σ(Xᵢ−X̄)²
- σ̂² uses divisor n (slightly biased) vs. S² with n−1 (unbiased)
- For large n, the difference is negligible
09 / 13
STAT 101
M05 · L01
More MLE results
Bernoulli & Poisson
- Bernoulli(p): p̂ = X̄ = k/n (sample proportion)
- Poisson(λ): λ̂ = X̄ (sample mean)
- Both match method of moments — not always the case
- Invariance: if θ̂ is MLE of θ, then g(θ̂) is MLE of g(θ)
10 / 13
STAT 101
M05 · L01
The fundamental tension
Bias–Variance Trade-off
MSE decomposition
\mathrm{MSE}(\hat{\theta}) = \mathrm{Var}(\hat{\theta}) + [\mathrm{Bias}(\hat{\theta})]^2
The insight
A biased estimator can beat an unbiased one if its variance is much smaller. Regularization exploits this.
11 / 13
STAT 101
Summary
Recap
What you learned
- Estimator (random) vs. estimate (fixed number)
- Unbiasedness, consistency, efficiency (CRLB)
- Method of moments: match sample & population moments
- MLE: maximize log-likelihood — consistent & asymptotically efficient
- MSE = Variance + Bias² — the core trade-off
13 / 13