STAT 101
M04 · L04
Module 4 · Lesson 4
The Central Limit Theorem
The most important theorem in statistics. It explains why averages are always normal — no matter the shape of the population.
01 / 13
STAT 101
M04 · L04
The setup
Sample Means are Random Variables
Draw n observations. Compute X̄. Do it again — you get a different X̄. The distribution of all possible X̄ values is the sampling distribution.
E[X̄]
μ (always)
Var(X̄)
σ² / n
02 / 13
STAT 101
M04 · L04
The formal statement
CLT: Convergence to Normal
For i.i.d. variables with mean μ and finite variance σ², as n → ∞:
Central Limit Theorem
Z_n = \dfrac{\bar{X}_n - \mu}{\sigma/\sqrt{n}} \xrightarrow{d} N(0,1)
03 / 13
STAT 101
M04 · L04
Any population, same result
Three Populations, One Shape
- Uniform[0,1]: already bell-shaped at n = 5
- Exponential(1): nearly normal by n = 30
- Bernoulli(0.3): normal by n = 30
The pattern
Skewed populations converge
slower than symmetric ones
slower than symmetric ones
04 / 13
STAT 101
M04 · L04
The practical question
How Large is Large Enough?
- Symmetric population: n = 5–10 suffices
- Moderately skewed: n ≥ 30 is the rule
- Heavily skewed / heavy-tailed: n = 100+
- Proportions: np ≥ 10 and n(1−p) ≥ 10
05 / 13
STAT 101
M04 · L04
Spread of the sampling distribution
Standard Error
How much do sample means vary from one sample to the next?
Standard Error
\text{SE}(\bar{X}) = \dfrac{\sigma}{\sqrt{n}}
Square-root law
To halve the SE: quadruple n
06 / 13
STAT 101
M04 · L04
The practical form
X̄ is Approximately Normal
For large n, we can write:
Approximate distribution
\bar{X} \doteq N\!\left(\mu,\,\tfrac{\sigma^2}{n}\right)
This justifies using z-scores and normal tables for sample means.
07 / 13
STAT 101
M04 · L04
A critical special case
CLT for Proportions
p̂ is just a sample mean of 0/1 variables — CLT applies directly:
CLT for proportions
\dfrac{\hat{p}-p}{\sqrt{p(1-p)/n}} \xrightarrow{d} N(0,1)
08 / 13
STAT 101
M04 · L04
CLT powers all of inference
Real Applications
- Confidence intervals: X̄ ± 1.96 × SE
- Z-tests and t-tests for means
- A/B testing and opinion polls
- Quality control: fraction defective
- Regression coefficient inference
09 / 13
STAT 101
M04 · L04
Why the bell curve is universal
Normal is an Attractor
- Heights = sum of genetic + environmental effects
- Measurement error = sum of tiny fluctuations
- Electronic noise = many uncorrelated micro-events
The reason
CLT: sums of independent
influences → Normal
influences → Normal
10 / 13
STAT 101
M04 · L04
What the CLT does NOT say
Common Misconceptions
- The raw data do NOT become normal — only the sample means do
- CLT fails for infinite-variance distributions (e.g., Cauchy)
- “n ≥ 30” is a heuristic, not a law — check skewness
- Large n does not fix biased samples
11 / 13
STAT 101
Summary
Recap
What you learned
The CLT: sample means converge to Normal for any population with finite variance. Standard error = σ/√n. CLT underpins all of classical inference.
Module 4 complete
Next: Module 5 — Estimation →
13 / 13