Reading
Stories Mode

Common Tests

~26 min read Lesson 2 of 4 in Module 6

A Toolkit of Tests

The previous lesson established the logic of hypothesis testing: state H₀ and H₁, compute a test statistic, find the p-value, and compare it to your significance threshold. But we left open a crucial question — which test do you use?

The answer depends on three things: what you are comparing (means, proportions, or associations), how many groups you have, and what you know about the population (e.g., whether the standard deviation is known). This lesson surveys the most commonly used tests: the Z-test, the one-sample and two-sample t-tests, the paired t-test, proportion tests, and the chi-squared test. We close with a discussion of effect size, which puts statistical significance in practical context.

Test Use When Key Statistic
Z-test (mean) One sample mean, σ known Z = (x̄ − μ₀) / (σ/√n)
One-sample t-test One sample mean, σ unknown t = (x̄ − μ₀) / (s/√n)
Two-sample t-test Two independent group means t = (x̄₁ − x̄₂) / SE
Paired t-test Two matched/related samples t = d̄ / (sₜ/√n)
One-proportion Z Single proportion vs. p₀ Z = (p̂ − p₀) / √(p₀(1−p₀)/n)
Two-proportion Z Two independent proportions Z = (p̂₁ − p̂₂) / SE
Chi-squared (χ²) Association between categorical variables χ² = Σ(O−E)²/E

Z-Test for Means (Known σ)

When the population standard deviation σ is known, the Z-test is exact even for small samples (assuming normality) and approximately valid for large samples via the Central Limit Theorem. In practice, σ is rarely known; the Z-test mainly appears in textbooks or in scenarios where historical data provide a reliable σ.

Z-Test Statistic
Z = \dfrac{\bar{x} - \mu_0}{\sigma / \sqrt{n}}
x̄ is the sample mean, μ₀ is the hypothesized population mean, σ is the known population standard deviation, and n is the sample size. The critical values for α = 0.05 (two-tailed) are ±1.96.
Example

A manufacturer claims bolts have a mean diameter of 10 mm with σ = 0.2 mm (from long production history). A sample of 40 bolts gives x̄ = 10.06 mm. Is there evidence the mean has shifted?

Z = (10.06 − 10) / (0.2 / √40) = 0.06 / 0.0316 ≈ 1.90. With |Z| = 1.90 < 1.96, we fail to reject H₀ at α = 0.05 (though p ≈ 0.057, close to borderline).

One-Sample t-Test (Unknown σ)

In the far more common case where σ is unknown, we estimate it with the sample standard deviation s and use a t-statistic. The t-distribution has heavier tails than the standard normal, reflecting the additional uncertainty from estimating σ. As the sample size grows, the t-distribution converges to the standard normal.

One-Sample t-Statistic
t = \dfrac{\bar{x} - \mu_0}{s / \sqrt{n}}, \quad df = n - 1
s is the sample standard deviation. The degrees of freedom are df = n − 1. With df = ∞, the t-distribution equals the standard normal.

Assumptions: The population is approximately normal (or n ≥ 30 by CLT). Independence of observations. No extreme outliers (which can distort s and violate normality).

The critical values come from a t-table or software. For n = 20 (df = 19), the two-tailed critical value at α = 0.05 is t* = 2.093, larger than 1.96 — the price of not knowing σ.

Two-Sample t-Test (Independent Groups)

When comparing the means of two independent groups — a treatment group vs. a control group, men vs. women, machine A vs. machine B — we use the two-sample t-test. H₀: μ₁ = μ₂.

The most common version assumes equal variances (σ₁² = σ₂²) and uses a pooled estimate of the common variance:

Pooled Variance
s_p^2 = \dfrac{(n_1 - 1)s_1^2 + (n_2 - 1)s_2^2}{n_1 + n_2 - 2}
The pooled variance s²_p is a weighted average of the two sample variances, with weights equal to the degrees of freedom contributed by each group.
Two-Sample t-Statistic
t = \dfrac{\bar{x}_1 - \bar{x}_2}{s_p\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}}, \quad df = n_1 + n_2 - 2
df = n₁ + n₂ − 2. If the equal-variance assumption is doubtful, use Welch's t-test, which uses separate variance estimates and an approximated df (Welch–Satterthwaite equation).
Equal vs. Unequal Variances

Levene's test or an F-test can check whether the two group variances are plausibly equal. However, many statisticians recommend Welch's t-test by default — it performs well even when variances are equal, and it is the safer choice when they are not.

Paired t-Test (Matched Samples)

When measurements come in natural pairs — before/after, left hand/right hand, patient on drug vs. placebo (crossover design) — the two samples are not independent. A paired t-test exploits this structure by working with the within-pair differences dᵢ = xᵢ₁ − xᵢ₂.

By analyzing the differences, we eliminate between-subject variability, giving the paired test more power than the two-sample t-test when pairs are strongly correlated.

Paired t-Statistic
t = \dfrac{\bar{d}}{s_d / \sqrt{n}}, \quad df = n - 1
d̄ is the mean of the n pairwise differences, s_d is their standard deviation. df = n − 1 (where n is the number of pairs). H₀: μ_d = 0.
Example: Blood Pressure Before and After Treatment

12 patients have blood pressure measured before and after a new medication. The mean difference is d̄ = −8.3 mmHg with sₜ = 5.1 mmHg. t = −8.3 / (5.1/√12) = −8.3 / 1.47 ≈ −5.65. With df = 11 and |t| = 5.65, p < 0.001 — strong evidence the medication reduces blood pressure.

Proportion Tests

When the outcome is categorical (success/failure, yes/no), we work with proportions rather than means. The Z approximation applies when the sample is large enough that the binomial distribution is well approximated by the normal.

One-Proportion Z-Test: Testing whether a population proportion equals a hypothesized value p₀.

One-Proportion Z-Statistic
Z = \dfrac{\hat{p} - p_0}{\sqrt{\dfrac{p_0(1 - p_0)}{n}}}
p̂ = x/n is the sample proportion. The rule of thumb for the normal approximation to be valid: np₀ ≥ 10 and n(1 − p₀) ≥ 10.

Two-Proportion Z-Test: Comparing proportions from two independent groups. The pooled proportion p̂ = (x₁ + x₂) / (n₁ + n₂) is used in the standard error under H₀: p₁ = p₂.

Two-Proportion Z-Statistic
Z = \dfrac{\hat{p}_1 - \hat{p}_2}{\sqrt{\bar{p}(1-\bar{p})\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}}
p̄ is the pooled sample proportion under H₀. This test is the backbone of A/B testing in industry, where one compares click-through rates, conversion rates, etc.

Chi-Squared Test for Independence

The chi-squared test answers: are two categorical variables independent? We arrange the data in a contingency table and compare the observed counts O with the expected counts E we would see if the variables were independent.

Expected Counts Under Independence
E_{ij} = \dfrac{(\text{row}_i \text{ total}) \times (\text{col}_j \text{ total})}{n}
Row total × column total ÷ grand total. Under independence, each cell's expected count is determined solely by the marginal frequencies.
Chi-Squared Statistic
\chi^2 = \sum_{\text{all cells}} \dfrac{(O - E)^2}{E}, \quad df = (r-1)(c-1)
The sum is over all cells. df = (r − 1)(c − 1), where r and c are the number of rows and columns. Large χ² means observed counts diverge substantially from independence.

Conditions: Expected count ≥ 5 in each cell (otherwise use Fisher's exact test). Random sampling and independence between observations.

Example: Drug Side-Effect vs. Treatment Group

In a trial, 200 patients receive drug A or B and report (or not) a side effect. Is side-effect incidence associated with drug type?

χ² summarizes how far the 2×2 table deviates from what you’d expect if there were no association. A large χ² with small p-value means the drugs differ in their side-effect profiles.

Effect Size: Beyond Statistical Significance

A p-value tells you whether an effect is detectable, not whether it is important. With a large enough sample, even a trivially small difference will be statistically significant. Effect size measures the magnitude of the effect in a standardized, interpretable unit.

Cohen’s d is the most widely used effect size for comparing two means. It expresses the difference in means as a multiple of the pooled standard deviation:

Cohen's d
d = \dfrac{\bar{x}_1 - \bar{x}_2}{s_p}
Conventional thresholds: |d| ≈ 0.2 (small), |d| ≈ 0.5 (medium), |d| ≈ 0.8 (large). These are rough guides — domain context always matters more.

Odds ratio (OR) is the natural effect size for binary outcomes and is widely used in epidemiology and clinical trials. The odds ratio compares the odds of an event in one group to the odds in another:

Odds Ratio
OR = \dfrac{p_1 / (1 - p_1)}{p_2 / (1 - p_2)} = \dfrac{ad}{bc}
OR = 1 means equal odds (no association). OR > 1 means higher odds in group 1. OR < 1 means lower odds. Unlike absolute risk difference, OR is invariant to the choice of reference group.

Always report both the p-value and the effect size. A result can be highly significant (tiny p) but practically meaningless (tiny d), or practically important (large d) but non-significant due to a small sample. Neither piece of information is complete without the other.

Key Takeaways
  • Z-test (known σ): Z = (x̄ − μ₀) / (σ/√n). Use when the population SD is known. Critical values from the standard normal.
  • One-sample t-test (unknown σ): t = (x̄ − μ₀) / (s/√n), df = n − 1. The workhorse for single-group mean tests.
  • Two-sample t-test: Compares two independent group means using a pooled (or Welch) variance. Prefer Welch’s version by default.
  • Paired t-test: Analyze differences dᵢ = xᵢ₁ − xᵢ₂ for matched pairs. More powerful than two-sample t when pairs are correlated.
  • Proportion Z-tests: One-proportion tests a single p̂ against p₀. Two-proportion tests compares p̂₁ vs. p̂₂ using a pooled SE.
  • Chi-squared test: χ² = Σ(O−E)²/E tests independence between two categorical variables. df = (r−1)(c−1). Requires E ≥ 5 per cell.
  • Effect size: Cohen’s d standardizes mean differences. Odds ratio compares binary event rates. Always report alongside p-values.
Previous Logic of Hypothesis Testing Module Overview Next ANOVA