Reading
Stories Mode

Confidence Intervals

~22 min read Lesson 2 of 3 in Module 5

From Point Estimate to Interval

A point estimate gives us a single best guess for an unknown parameter — but it gives us no sense of how reliable that guess is. Does it come from 10 observations or 10,000? Is the underlying distribution nearly normal or wildly skewed? A point estimate alone cannot answer these questions.

A confidence interval (CI) extends the point estimate into a range of plausible values, communicating our uncertainty about the parameter. Instead of reporting “μ ≈ 4.7,” we report “we are 95% confident that μ lies between 4.1 and 5.3.” The width of the interval encodes information about sample size, variability, and the chosen confidence level.

This lesson builds confidence intervals for the three most important cases: the population mean when the population variance is known (z-interval), the population mean when the variance is unknown (t-interval), and a population proportion. We also address the most common misconception in all of statistics — what “confidence” actually means — and show how to choose a sample size to achieve a desired interval width.

What a Confidence Interval Is NOT

Before defining what a confidence interval is, let us dispel the most pervasive misinterpretation in statistics. Many students, and even many scientists, believe that a 95% confidence interval means: “there is a 95% probability that the true parameter θ lies inside this interval.”

This is wrong. The true parameter θ is a fixed (though unknown) number — it does not have a probability distribution. Once we compute an interval from data, that specific interval either contains θ or it does not. There is no probability involved.

The Correct Interpretation

A 95% confidence interval is a procedure, not a statement about a single interval. If we repeated our experiment many times and computed a 95% CI each time, then approximately 95% of those intervals would contain the true parameter.

For any single computed interval, we say we are “95% confident” — meaning the interval was produced by a method that succeeds 95% of the time — not that there is a 95% probability the parameter is inside.

CI for the Mean — σ Known (z-interval)

The simplest confidence interval assumes we know the population standard deviation σ (rare in practice, but a useful starting point). By the Central Limit Theorem, the sample mean X̄ is approximately normally distributed for large n:

The standardized statistic Z = (X̄ − μ) / (σ/√n) follows a standard normal distribution. Inverting this pivot gives the confidence interval:

z-Confidence Interval for μ
\bar{x} \pm z_{\alpha/2}\,\dfrac{\sigma}{\sqrt{n}}
x̄ is the sample mean, σ is the known population standard deviation, n is sample size, and z_{α/2} is the critical value from the standard normal. For 95% CI: z_{0.025} = 1.96.

The term zα/2 ⋅ σ/√n is called the margin of error (E). It quantifies how far we expect our sample mean to be from the true mean. Larger samples produce narrower intervals; higher confidence levels require larger zα/2 and thus wider intervals.

Common Critical Values

90% CI: α = 0.10, zα/2 = z0.05 = 1.645

95% CI: α = 0.05, zα/2 = z0.025 = 1.960

99% CI: α = 0.01, zα/2 = z0.005 = 2.576

The t-Distribution and Degrees of Freedom

In practice, we almost never know σ. When we replace σ with the sample standard deviation S, the standardized statistic no longer follows a standard normal distribution — it follows a t-distribution with n − 1 degrees of freedom:

t-Statistic
t = \dfrac{\bar{X} - \mu}{S/\sqrt{n}} \sim t_{n-1}
When σ is replaced by the sample standard deviation S, the pivot follows a t-distribution with ν = n − 1 degrees of freedom, not a standard normal.

The t-distribution looks like the standard normal but has heavier tails, reflecting the additional uncertainty from estimating σ. As n → ∞, the t-distribution converges to the standard normal. The “degrees of freedom” parameter ν = n − 1 counts the number of independent pieces of information used to estimate σ (we lose one because the mean is estimated first).

The t-distribution was derived by William Gosset in 1908, publishing under the pseudonym “Student” because his employer (Guinness Brewery) required anonymity. Hence the name Student’s t-distribution. His contribution was recognizing that small-sample inference required a fundamentally different distribution than the normal.

CI for the Mean — σ Unknown (t-interval)

With S replacing σ and t critical values replacing z critical values, the confidence interval for μ when σ is unknown is:

t-Confidence Interval for μ
\bar{x} \pm t_{\alpha/2,\,n-1}\,\dfrac{s}{\sqrt{n}}
s is the sample standard deviation, t_{α/2, n−1} is the t critical value with n−1 degrees of freedom. This is valid when the population is approximately normal or n is large enough for the CLT.

This is the t-interval, and it is by far the most commonly used confidence interval in practice. Because tα/2, n−1 > zα/2 for any finite n, the t-interval is always wider than the z-interval — reflecting the additional uncertainty from not knowing σ.

Example: Suppose n = 16 observations from a normal population yield x̄ = 12.3 and s = 2.4. For a 95% CI, we need t0.025, 15 = 2.131. The interval is: 12.3 ± 2.131 ⋅ 2.4/√16 = 12.3 ± 1.279 = (11.02, 13.58).

When is the t-interval valid?

Population is normal: The t-interval is exact for any sample size n.

Large sample (n ≥ 30): The CLT makes the t-interval approximately valid even for non-normal populations. For moderately skewed populations, n ≥ 50 is safer.

Outliers: The t-interval is sensitive to outliers when n is small. Check for outliers with a boxplot or normal probability plot first.

CI for a Proportion

When the parameter of interest is a population proportion p (e.g., the fraction of voters who support a candidate, or the defect rate in manufacturing), the natural estimator is the sample proportion p̂ = X/n, where X is the number of “successes” in n trials.

By the CLT, p̂ is approximately normal for large n, with mean p and standard error √(p(1−p)/n). Replacing p with p̂ in the standard error gives the Wald confidence interval for proportions:

CI for a Proportion
\hat{p} \pm z_{\alpha/2}\sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}
p̂ = X/n is the sample proportion, n is sample size, and z_{α/2} is the standard normal critical value. Valid when np̂ ≥ 10 and n(1−p̂) ≥ 10.

Example: In a survey of n = 400 voters, X = 220 say they will vote for candidate A. Then p̂ = 220/400 = 0.55. A 95% CI is: 0.55 ± 1.96 ⋅ √(0.55 ⋅ 0.45 / 400) = 0.55 ± 0.0487 = (0.501, 0.599).

The Wald interval has known deficiencies for extreme proportions (p near 0 or 1) or small samples. The Wilson interval (also called score interval) performs better in these cases and is preferred in modern practice, but the Wald interval remains common for introductory contexts.

Interpreting Confidence Levels

Confidence levels are set by the analyst before collecting data. The three most common choices are 90%, 95%, and 99%. What does choosing 95% mean in practice?

Think of the confidence interval procedure as a machine. Feed it data from repeated experiments; it produces intervals. For a 95% procedure, 95 out of every 100 intervals produced by this machine will contain the true parameter. The other 5 will miss.

Higher confidence levels come at a cost: wider intervals. Intuitively, if you want to be more certain you have captured the true value, you must expand your search range. The relationship is deterministic: doubling the confidence level from 90% to 99% increases the critical value from 1.645 to 2.576, widening the interval by a factor of about 1.57.

The Three Levers of Interval Width

Increase n: Standard error shrinks as 1/√n — quadrupling n halves the margin of error.

Decrease confidence level: Lower α means smaller z or t critical value — but you accept a higher miss rate.

Decrease σ (or s): More precise measurements or a more homogeneous population reduce the interval width. Usually out of our control.

Sample Size Determination

Before collecting data, we often want to know: how large a sample do I need to achieve a desired margin of error E at confidence level (1 − α)?

For estimating a mean with known σ, set the margin of error equal to E and solve for n:

Sample Size for a Mean
n = \left(\dfrac{z_{\alpha/2}\,\sigma}{E}\right)^{\!2}
Always round up to the next integer. If σ is unknown, use a preliminary estimate from a pilot study, or a conservative upper bound.

For estimating a proportion, a preliminary estimate p̂ is needed. In the absence of prior information, use p̂ = 0.5, which maximizes the variance p(1−p) and gives the most conservative (largest) sample size:

Sample Size for a Proportion
n = \left(\dfrac{z_{\alpha/2}}{E}\right)^{\!2}\hat{p}(1-\hat{p})
With p̂ = 0.5 (maximum variance), for E = 0.03 (3%) and 95% confidence: n = (1.96/0.03)² × 0.25 ≈ 1068. This is why national polls typically survey ~1000 people for a ±3% margin.

Note on the t-interval: For sample size planning with unknown σ, the formula above uses zα/2 as an approximation. For small n, the t critical value depends on n itself (through degrees of freedom), requiring an iterative solution. Software handles this automatically.

Key Takeaways
  • What a CI is: a procedure that produces intervals; 95% of intervals from such a procedure contain the true parameter. NOT the probability that this specific interval contains θ.
  • z-interval (σ known): x̄ ± zα/2 ⋅ σ/√n. Requires known population variance; rarely applies in practice.
  • t-distribution: when σ is replaced by S, the pivot follows tn−1 with heavier tails than the normal, converging to normal as n → ∞.
  • t-interval (σ unknown): x̄ ± tα/2, n−1 ⋅ s/√n. The standard interval in practice; valid when the population is approximately normal or n is large.
  • Proportion CI: p̂ ± zα/2 ⋅ √(p̂(1−p̂)/n). Valid when np̂ ≥ 10 and n(1−p̂) ≥ 10.
  • Confidence level trade-off: higher confidence ⇒ wider interval. Interval width also shrinks with √n; quadrupling n halves the margin of error.
  • Sample size planning: set E = zα/2 ⋅ σ/√n and solve for n; for proportions use p̂ = 0.5 as a conservative default.
Previous Point Estimation Module Overview Next Bootstrap Methods