STAT 101
M04 · L04
Module 4 · Lesson 4

The Central Limit Theorem

The most important theorem in statistics. It explains why averages are always normal — no matter the shape of the population.

01 / 13
STAT 101
M04 · L04
The setup

Sample Means are Random Variables

Draw n observations. Compute X̄. Do it again — you get a different X̄. The distribution of all possible X̄ values is the sampling distribution.

E[X̄]
μ (always)
Var(X̄)
σ² / n
02 / 13
STAT 101
M04 · L04
The formal statement

CLT: Convergence to Normal

For i.i.d. variables with mean μ and finite variance σ², as n → ∞:

Central Limit Theorem
Z_n = \dfrac{\bar{X}_n - \mu}{\sigma/\sqrt{n}} \xrightarrow{d} N(0,1)
03 / 13
STAT 101
M04 · L04
Any population, same result

Three Populations, One Shape

  • Uniform[0,1]: already bell-shaped at n = 5
  • Exponential(1): nearly normal by n = 30
  • Bernoulli(0.3): normal by n = 30
The pattern
Skewed populations converge
slower than symmetric ones
04 / 13
STAT 101
M04 · L04
The practical question

How Large is Large Enough?

  • Symmetric population: n = 5–10 suffices
  • Moderately skewed: n ≥ 30 is the rule
  • Heavily skewed / heavy-tailed: n = 100+
  • Proportions: np ≥ 10 and n(1−p) ≥ 10
05 / 13
STAT 101
M04 · L04
Spread of the sampling distribution

Standard Error

How much do sample means vary from one sample to the next?

Standard Error
\text{SE}(\bar{X}) = \dfrac{\sigma}{\sqrt{n}}
Square-root law
To halve the SE: quadruple n
06 / 13
STAT 101
M04 · L04
The practical form

X̄ is Approximately Normal

For large n, we can write:

Approximate distribution
\bar{X} \doteq N\!\left(\mu,\,\tfrac{\sigma^2}{n}\right)

This justifies using z-scores and normal tables for sample means.

07 / 13
STAT 101
M04 · L04
A critical special case

CLT for Proportions

p̂ is just a sample mean of 0/1 variables — CLT applies directly:

CLT for proportions
\dfrac{\hat{p}-p}{\sqrt{p(1-p)/n}} \xrightarrow{d} N(0,1)
08 / 13
STAT 101
M04 · L04
CLT powers all of inference

Real Applications

  • Confidence intervals: X̄ ± 1.96 × SE
  • Z-tests and t-tests for means
  • A/B testing and opinion polls
  • Quality control: fraction defective
  • Regression coefficient inference
09 / 13
STAT 101
M04 · L04
Why the bell curve is universal

Normal is an Attractor

  • Heights = sum of genetic + environmental effects
  • Measurement error = sum of tiny fluctuations
  • Electronic noise = many uncorrelated micro-events
The reason
CLT: sums of independent
influences → Normal
10 / 13
STAT 101
M04 · L04
What the CLT does NOT say

Common Misconceptions

  • The raw data do NOT become normal — only the sample means do
  • CLT fails for infinite-variance distributions (e.g., Cauchy)
  • “n ≥ 30” is a heuristic, not a law — check skewness
  • Large n does not fix biased samples
11 / 13
STAT 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
STAT 101
Summary
Recap

What you learned

The CLT: sample means converge to Normal for any population with finite variance. Standard error = σ/√n. CLT underpins all of classical inference.

Module 4 complete
Next: Module 5 — Estimation →
13 / 13