A Toolkit of Tests
The previous lesson established the logic of hypothesis testing: state H₀ and H₁, compute a test statistic, find the p-value, and compare it to your significance threshold. But we left open a crucial question — which test do you use?
The answer depends on three things: what you are comparing (means, proportions, or associations), how many groups you have, and what you know about the population (e.g., whether the standard deviation is known). This lesson surveys the most commonly used tests: the Z-test, the one-sample and two-sample t-tests, the paired t-test, proportion tests, and the chi-squared test. We close with a discussion of effect size, which puts statistical significance in practical context.
| Test | Use When | Key Statistic |
|---|---|---|
| Z-test (mean) | One sample mean, σ known | Z = (x̄ − μ₀) / (σ/√n) |
| One-sample t-test | One sample mean, σ unknown | t = (x̄ − μ₀) / (s/√n) |
| Two-sample t-test | Two independent group means | t = (x̄₁ − x̄₂) / SE |
| Paired t-test | Two matched/related samples | t = d̄ / (sₜ/√n) |
| One-proportion Z | Single proportion vs. p₀ | Z = (p̂ − p₀) / √(p₀(1−p₀)/n) |
| Two-proportion Z | Two independent proportions | Z = (p̂₁ − p̂₂) / SE |
| Chi-squared (χ²) | Association between categorical variables | χ² = Σ(O−E)²/E |
Z-Test for Means (Known σ)
When the population standard deviation σ is known, the Z-test is exact even for small samples (assuming normality) and approximately valid for large samples via the Central Limit Theorem. In practice, σ is rarely known; the Z-test mainly appears in textbooks or in scenarios where historical data provide a reliable σ.
A manufacturer claims bolts have a mean diameter of 10 mm with σ = 0.2 mm (from long production history). A sample of 40 bolts gives x̄ = 10.06 mm. Is there evidence the mean has shifted?
Z = (10.06 − 10) / (0.2 / √40) = 0.06 / 0.0316 ≈ 1.90. With |Z| = 1.90 < 1.96, we fail to reject H₀ at α = 0.05 (though p ≈ 0.057, close to borderline).
One-Sample t-Test (Unknown σ)
In the far more common case where σ is unknown, we estimate it with the sample standard deviation s and use a t-statistic. The t-distribution has heavier tails than the standard normal, reflecting the additional uncertainty from estimating σ. As the sample size grows, the t-distribution converges to the standard normal.
Assumptions: The population is approximately normal (or n ≥ 30 by CLT). Independence of observations. No extreme outliers (which can distort s and violate normality).
The critical values come from a t-table or software. For n = 20 (df = 19), the two-tailed critical value at α = 0.05 is t* = 2.093, larger than 1.96 — the price of not knowing σ.
Two-Sample t-Test (Independent Groups)
When comparing the means of two independent groups — a treatment group vs. a control group, men vs. women, machine A vs. machine B — we use the two-sample t-test. H₀: μ₁ = μ₂.
The most common version assumes equal variances (σ₁² = σ₂²) and uses a pooled estimate of the common variance:
Levene's test or an F-test can check whether the two group variances are plausibly equal. However, many statisticians recommend Welch's t-test by default — it performs well even when variances are equal, and it is the safer choice when they are not.
Paired t-Test (Matched Samples)
When measurements come in natural pairs — before/after, left hand/right hand, patient on drug vs. placebo (crossover design) — the two samples are not independent. A paired t-test exploits this structure by working with the within-pair differences dᵢ = xᵢ₁ − xᵢ₂.
By analyzing the differences, we eliminate between-subject variability, giving the paired test more power than the two-sample t-test when pairs are strongly correlated.
12 patients have blood pressure measured before and after a new medication. The mean difference is d̄ = −8.3 mmHg with sₜ = 5.1 mmHg. t = −8.3 / (5.1/√12) = −8.3 / 1.47 ≈ −5.65. With df = 11 and |t| = 5.65, p < 0.001 — strong evidence the medication reduces blood pressure.
Proportion Tests
When the outcome is categorical (success/failure, yes/no), we work with proportions rather than means. The Z approximation applies when the sample is large enough that the binomial distribution is well approximated by the normal.
One-Proportion Z-Test: Testing whether a population proportion equals a hypothesized value p₀.
Two-Proportion Z-Test: Comparing proportions from two independent groups. The pooled proportion p̂ = (x₁ + x₂) / (n₁ + n₂) is used in the standard error under H₀: p₁ = p₂.
Chi-Squared Test for Independence
The chi-squared test answers: are two categorical variables independent? We arrange the data in a contingency table and compare the observed counts O with the expected counts E we would see if the variables were independent.
Conditions: Expected count ≥ 5 in each cell (otherwise use Fisher's exact test). Random sampling and independence between observations.
In a trial, 200 patients receive drug A or B and report (or not) a side effect. Is side-effect incidence associated with drug type?
χ² summarizes how far the 2×2 table deviates from what you’d expect if there were no association. A large χ² with small p-value means the drugs differ in their side-effect profiles.
Effect Size: Beyond Statistical Significance
A p-value tells you whether an effect is detectable, not whether it is important. With a large enough sample, even a trivially small difference will be statistically significant. Effect size measures the magnitude of the effect in a standardized, interpretable unit.
Cohen’s d is the most widely used effect size for comparing two means. It expresses the difference in means as a multiple of the pooled standard deviation:
Odds ratio (OR) is the natural effect size for binary outcomes and is widely used in epidemiology and clinical trials. The odds ratio compares the odds of an event in one group to the odds in another:
Always report both the p-value and the effect size. A result can be highly significant (tiny p) but practically meaningless (tiny d), or practically important (large d) but non-significant due to a small sample. Neither piece of information is complete without the other.
- Z-test (known σ): Z = (x̄ − μ₀) / (σ/√n). Use when the population SD is known. Critical values from the standard normal.
- One-sample t-test (unknown σ): t = (x̄ − μ₀) / (s/√n), df = n − 1. The workhorse for single-group mean tests.
- Two-sample t-test: Compares two independent group means using a pooled (or Welch) variance. Prefer Welch’s version by default.
- Paired t-test: Analyze differences dᵢ = xᵢ₁ − xᵢ₂ for matched pairs. More powerful than two-sample t when pairs are correlated.
- Proportion Z-tests: One-proportion tests a single p̂ against p₀. Two-proportion tests compares p̂₁ vs. p̂₂ using a pooled SE.
- Chi-squared test: χ² = Σ(O−E)²/E tests independence between two categorical variables. df = (r−1)(c−1). Requires E ≥ 5 per cell.
- Effect size: Cohen’s d standardizes mean differences. Odds ratio compares binary event rates. Always report alongside p-values.