Reading
Stories Mode

ANOVA

~28 min read Lesson 3 of 4 in Module 6

The Problem with Multiple t-Tests

Suppose you are comparing the mean response times of four different web server configurations. Your first instinct might be to run a t-test on every pair: A vs. B, A vs. C, A vs. D, B vs. C, B vs. D, and C vs. D — six tests in total. But doing so at α = 0.05 per test inflates your overall chance of a false positive far beyond 5%.

If all six tests are independent (they are not, since they share group means), the probability of at least one false rejection is 1 − (0.95)⁶ ≈ 26%. With more groups, this familywise error rate balloons rapidly. Analysis of Variance (ANOVA) solves this by testing all groups simultaneously with a single test statistic, keeping the Type I error at the nominal α.

Key Insight

ANOVA does not test whether all means are equal to some constant. It tests H₀: μ₁ = μ₂ = … = μₖ against H₁: at least one mean differs. A significant ANOVA tells you something is different; post-hoc tests tell you what.

One-Way ANOVA: Partitioning Variance

The central idea of ANOVA is partitioning the total variability in the data into two parts: variability between groups (signal) and variability within groups (noise). If the between-group variability is large relative to the within-group variability, we have evidence that the groups differ.

Formally, with k groups each of size nᵢ and grand mean x̄, we decompose the Total Sum of Squares:

SS Decomposition
SS_T = SS_B + SS_W = \sum_{i=1}^{k}\sum_{j=1}^{n_i}(x_{ij} - \bar{x})^2
SS_T = SS_B + SS_W. The total deviation of each observation from the grand mean equals the group-mean deviation (between) plus the within-group deviation.
Between-Group SS
SS_B = \sum_{i=1}^{k} n_i (\bar{x}_i - \bar{x})^2
Measures how far each group mean x̄ᵢ is from the grand mean x̄, weighted by group size nᵢ. Large SS_B means groups have very different means.
Within-Group SS
SS_W = \sum_{i=1}^{k}\sum_{j=1}^{n_i}(x_{ij} - \bar{x}_i)^2
Measures the spread of observations around their own group mean. This is the "noise" — variability that cannot be explained by group membership.

The F-Statistic

Dividing each sum of squares by its degrees of freedom gives a Mean Square. The ratio of the two Mean Squares is the F-statistic:

F-Statistic
F = \dfrac{MS_B}{MS_W} = \dfrac{SS_B/(k-1)}{SS_W/(N-k)}
MS_B = SS_B / (k − 1), MS_W = SS_W / (N − k), where N = total observations. Under H₀, F follows an F-distribution with df₁ = k − 1 and df₂ = N − k.

The ANOVA table organizes these calculations in a compact form:

Source SS df MS F
Between Groups SSB k − 1 SSB / (k−1) MSB / MSW
Within Groups SSW N − k SSW / (N−k) —
Total SSₜ N − 1 — —

Under H₀ (all means equal), both MSB and MSW estimate the same population variance σ², so F ≈ 1. Under H₁ (some means differ), MSB overestimates σ² while MSW remains an unbiased estimate, pushing F > 1. We reject H₀ when F exceeds the critical value F*(k−1, N−k) at level α.

Assumptions of One-Way ANOVA

Like the t-test, ANOVA rests on three core assumptions that must be checked before trusting results:

1. Normality: Within each group, the observations should be approximately normally distributed. For large samples, the Central Limit Theorem makes ANOVA reasonably robust to non-normality. Check with QQ plots or a Shapiro-Wilk test per group.

2. Homoscedasticity (Equal Variances): All groups should have the same population variance σ². This is the most critical assumption. Check with Levene’s test or by plotting group-level standard deviations. As a rough guide, the largest s should be no more than twice the smallest s.

3. Independence: Observations must be independent both within and across groups. Violation of independence (e.g., repeated measures, clustering) requires a different model entirely.

What If Assumptions Fail?

If normality is severely violated: use Kruskal-Wallis (covered later in this lesson). If variances are unequal: use Welch’s ANOVA (extends Welch’s t-test idea to k groups). If observations are not independent: use repeated measures ANOVA or a mixed-effects model.

Post-Hoc Tests: Finding Which Groups Differ

A significant F-test confirms that not all group means are equal, but it does not say which pairs differ. Post-hoc tests answer this follow-up question while controlling the familywise error rate.

Tukey’s HSD (Honestly Significant Difference) is the most widely used post-hoc test. It computes a critical difference: two group means are declared significantly different if their absolute difference exceeds HSD:

Tukey HSD
HSD = q^* \sqrt{\dfrac{MS_W}{n}}
q* is the critical value from the studentized range distribution with parameters k (number of groups) and df_W = N − k. n is the group size (assumes equal group sizes; Tukey-Kramer extends to unequal sizes).

Bonferroni correction is simpler: divide α by the number of comparisons m = k(k−1)/2 and use this adjusted threshold for each individual t-test. It is more conservative than Tukey’s HSD but more flexible — it applies to any set of planned comparisons, not just pairwise.

Choosing a Post-Hoc Test

Tukey’s HSD is best for all pairwise comparisons with equal group sizes — it maximizes power while controlling familywise error. Bonferroni is preferred when you have a smaller, pre-specified set of comparisons or when Tukey’s is not available in your software. Scheffé’s test is for arbitrary contrasts (not just pairwise), but at the cost of less power.

Two-Way ANOVA

One-way ANOVA studies a single factor (independent variable). Two-way ANOVA extends this to two factors simultaneously, asking three questions at once: Does Factor A have an effect? Does Factor B have an effect? Do A and B interact?

With two factors A (a levels) and B (b levels), the model partitions the total variance as:

Two-Way SS Decomposition
SS_T = SS_A + SS_B + SS_{AB} + SS_E
SS_A, SS_B are the main effect sums of squares; SS_AB is the interaction sum of squares; SS_E is the residual (within-cell) error. Each term has its own F-test.

The three F-tests are:

Test F-Ratio df numerator df denominator
Main effect A MSₐ / MSE a − 1 ab(n−1)
Main effect B MSₑ / MSE b − 1 ab(n−1)
Interaction AB MSₐₑ / MSE (a−1)(b−1) ab(n−1)

Interaction Effects

An interaction occurs when the effect of one factor depends on the level of another. This is the most important and often most surprising output of two-way ANOVA.

Imagine testing two teaching methods (lecture vs. interactive) on two student groups (beginners vs. advanced). If interactive teaching improves beginners but not advanced students, the method effect depends on skill level — that is an interaction.

Visualizing Interactions

Plot the group means as lines: put one factor on the x-axis and draw a separate line for each level of the other factor. Non-parallel lines suggest an interaction — the steeper the crossing or divergence, the stronger the interaction. Parallel lines indicate no interaction: each factor’s effect is consistent across levels of the other.

Critical rule: When a significant interaction is present, main effects should not be interpreted in isolation. The interaction tells you the main effect story changes depending on which level of the other factor you are in.

Kruskal-Wallis: The Non-Parametric Alternative

When the normality assumption is clearly violated and the sample size is too small for the CLT to rescue you, the Kruskal-Wallis test is the rank-based alternative to one-way ANOVA. It requires no distributional assumption beyond continuity.

The procedure: rank all N observations jointly from 1 to N, then compute the H-statistic based on how much the average rank within each group deviates from the overall average rank (N+1)/2:

Kruskal-Wallis H
H = \dfrac{12}{N(N+1)} \sum_{i=1}^{k} \dfrac{R_i^2}{n_i} - 3(N+1)
Rᵢ is the sum of ranks in group i, nᵢ is the group size, and N is the total sample size. Under H₀, H approximately follows a chi-squared distribution with k − 1 degrees of freedom.

H is large when the rank sums Rᵢ differ substantially across groups, indicating the groups come from different distributions. The p-value comes from the χ²(k−1) distribution (exact for small n, approximate for large n).

The trade-off: Kruskal-Wallis has about 95% of the power of one-way ANOVA when normality holds, and greater power when it doesn’t. The cost is that it tests for a general difference in distributions, not specifically in means. Post-hoc comparisons after a significant Kruskal-Wallis use Dunn’s test with a Bonferroni or Benjamini-Hochberg correction.

Key Takeaways
  • Multiple t-tests inflate error: Running m pairwise tests at α each drives the familywise error far above α. ANOVA keeps it controlled with a single F-test.
  • Variance partitioning: SSₜ = SSB + SSW. The F-statistic = MSB / MSW measures signal-to-noise. Large F ⇒ evidence against H₀.
  • Assumptions: Normality per group, equal variances (Levene’s test), independence. Violations call for Welch’s ANOVA or Kruskal-Wallis.
  • Post-hoc tests: Significant F only says “something differs.” Tukey’s HSD (pairwise, equal n) or Bonferroni (any pre-planned comparisons) identifies which pairs.
  • Two-way ANOVA tests two main effects and their interaction simultaneously. Check the interaction first — it can reverse or mask main effects.
  • Interaction: Non-parallel lines in an interaction plot. When significant, do not interpret main effects alone.
  • Kruskal-Wallis: Rank-based alternative to one-way ANOVA. H ∼ χ²(k−1) under H₀. Use when normality is violated and n is small.
Previous Common Tests Module Overview Next Common Pitfalls