STAT 101
M05 · L01
Module 5 · Lesson 1

Point Estimation

We observe data. We want to know the underlying parameter. Point estimation gives us a single best guess — and the tools to know how good that guess is.

01 / 13
STAT 101
M05 · L01
Precision of language

Estimator vs. Estimate

Estimator
g(X₁,…,Xₙ)
random variable
Estimate
g(x₁,…,xₙ)
specific number

Upper case = random variable. Lower case = observed value. Properties like bias and variance belong to the estimator, not any single estimate.

02 / 13
STAT 101
M05 · L01
Desirable property #1

Unbiasedness

On average, over all possible samples, the estimator hits the true target:

Unbiasedness
E[\hat{\theta}] = \theta
Example
X̄ is unbiased for μ.
S² with n−1 is unbiased for σ².
03 / 13
STAT 101
M05 · L01
Desirable property #2

Consistency

More data = closer to the truth. Formally, θ̂ₙ converges in probability to θ:

Consistency
\lim_{n\to\infty} P(|\hat{\theta}_n - \theta| > \varepsilon) = 0
04 / 13
STAT 101
M05 · L01
Desirable property #3

Efficiency & CRLB

  • Among all unbiased estimators, prefer the lowest variance
  • The Cramér-Rao bound sets the minimum achievable variance
  • An estimator hitting the bound is called UMVUE
  • The MLE is asymptotically efficient
05 / 13
STAT 101
M05 · L01
Estimation framework #1

Method of Moments

Set sample moments equal to population moments, then solve:

MOM
E[X^k] = \tfrac{1}{n}\sum_{i=1}^n X_i^k
Normal example
μ̂ = X̄, σ̂² = (1/n)Σ(Xᵢ−X̄)²
06 / 13
STAT 101
M05 · L01
Estimation framework #2

Maximum Likelihood

Choose θ to make the observed data as probable as possible:

Likelihood
L(\theta) = \prod_{i=1}^{n} f(x_i;\,\theta)

In practice, maximize the log-likelihood ℓ(θ) = Σ log f(xᵢ; θ).

07 / 13
STAT 101
M05 · L01
Solving for the MLE

The Score Equation

Differentiate log-likelihood, set to zero:

Score equation
\dfrac{\partial\,\ell(\theta)}{\partial\theta} = 0 \implies \hat{\theta}_{\mathrm{MLE}}
Why log?
Products → sums. Numerically stable. Same maximizer.
08 / 13
STAT 101
M05 · L01
MLE in practice

MLE for the Normal

  • Maximize over μ → μ̂ = X̄ (sample mean)
  • Maximize over σ² → σ̂² = (1/n)Σ(Xᵢ−X̄)²
  • σ̂² uses divisor n (slightly biased) vs. S² with n−1 (unbiased)
  • For large n, the difference is negligible
09 / 13
STAT 101
M05 · L01
More MLE results

Bernoulli & Poisson

  • Bernoulli(p): p̂ = X̄ = k/n (sample proportion)
  • Poisson(λ): λ̂ = X̄ (sample mean)
  • Both match method of moments — not always the case
  • Invariance: if θ̂ is MLE of θ, then g(θ̂) is MLE of g(θ)
10 / 13
STAT 101
M05 · L01
The fundamental tension

Bias–Variance Trade-off

MSE decomposition
\mathrm{MSE}(\hat{\theta}) = \mathrm{Var}(\hat{\theta}) + [\mathrm{Bias}(\hat{\theta})]^2
The insight
A biased estimator can beat an unbiased one if its variance is much smaller. Regularization exploits this.
11 / 13
STAT 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
STAT 101
Summary
Recap

What you learned

  • Estimator (random) vs. estimate (fixed number)
  • Unbiasedness, consistency, efficiency (CRLB)
  • Method of moments: match sample & population moments
  • MLE: maximize log-likelihood — consistent & asymptotically efficient
  • MSE = Variance + Bias² — the core trade-off
13 / 13