Reading
Stories Mode

Why Non-Parametric?

~30 min read Lesson 1 of 3 in Module 8

When the Assumptions Break

Almost every test so far — the t-test, ANOVA, the F-test, the confidence intervals of Module 5 — rests on an assumption about the shape of the data, usually that it is normally distributed (or that the sample is large enough for the Central Limit Theorem to make the sampling distribution normal). These are parametric methods: they assume the data comes from a family described by a few parameters, and they estimate those parameters.

That assumption is a convenience, not a law of nature, and real data breaks it constantly. Non-parametric methods make no assumption about the distribution’s shape — they are “distribution-free.” This lesson is about when you need them and what you pay for the freedom; the next two show the specific tests and regression methods.

Four Situations That Break Parametric Tests

Reach for a non-parametric method when any of these holds:

The Core Idea: Ranks, Not Values

Almost every non-parametric test shares one trick: replace the raw values with their ranks, then work with those.

The Rank Transform
R_i = \text{rank of } x_i \text{ among } x_1, x_2, \dots, x_n
Sort the data and replace each value by its position. The smallest becomes 1, the next 2, and so on — the actual magnitudes are discarded.

Ranks are why these methods are distribution-free. Once you have thrown away everything except order, the exact shape of the original distribution no longer matters — the ranks of n numbers are always 1 through n, whatever the numbers were. That is also why a wild outlier loses its power to distort: it contributes one rank, not one enormous number.

The Trade-off: Power for Robustness

Nothing is free. When the parametric assumptions do hold, a non-parametric test has slightly less power — a somewhat higher chance of missing a real effect — because throwing away the magnitudes throws away information. The standard measure is the asymptotic relative efficiency (ARE):

The Cost, Quantified
\text{ARE}(\text{Wilcoxon}, t) = \dfrac{3}{\pi} \approx 0.955
On truly normal data the Wilcoxon test needs about 1/0.955 ≈ 1.05× the sample size of the t-test to match its power — a 5% penalty. On non-normal data it is frequently more powerful than the t-test.

A 5% loss when the assumptions hold, against near-total robustness when they do not, is usually a bargain. The catch that matters more in practice is interpretation: a rank test answers a question about medians or stochastic ordering, not about means, so you must state the conclusion in those terms.

What’s Ahead — and Already Here

Lesson 8.2 covers the specific tests — Wilcoxon, Mann-Whitney, Friedman, Kolmogorov-Smirnov, and permutation tests — and Lesson 8.3 covers non-parametric regression. Two pieces of this module are already taught elsewhere in the course and will be cross-referenced rather than repeated: the Kruskal-Wallis test is in Module 6, Lesson 3, and kernel density estimation is in Module 2, Lesson 1.

Key Takeaways
  • Parametric methods assume a distribution (usually normal); non-parametric methods are distribution-free.
  • Reach for non-parametric methods with small samples, ordinal data, heavy tails/outliers, or strong skew — wherever normality is untrue or uncheckable.
  • The shared trick is the rank transform: replace values by their order, which discards magnitudes and makes the test distribution-free.
  • Ranks make the methods robust to outliers — an extreme value is just the biggest rank.
  • The cost is a small power penalty (~5% for Wilcoxon vs. the t-test on normal data), often reversed on non-normal data.
  • Rank tests answer questions about medians / stochastic ordering, not means — state the conclusion accordingly.
Previous Beyond Linear Regression Module Overview Next Non-Parametric Tests