Reading
Stories Mode

Non-Parametric Tests

~30 min read Lesson 2 of 3 in Module 8

A Test for Every Design

Every parametric test has a rank-based counterpart that asks a similar question without assuming normality. The design of your study — two groups or many, paired or independent — picks the test, exactly as it does in the parametric world:

Two Independent Groups: Mann-Whitney U

The Mann-Whitney U test (equivalent to the Wilcoxon rank-sum test) asks whether one group tends to produce larger values than the other. Pool both samples, rank every observation together, then sum the ranks belonging to group 1:

The U Statistic
U = n_1 n_2 + \dfrac{n_1(n_1+1)}{2} - R_1
R₁ is the sum of ranks in group 1. U counts how often a value from group 1 exceeds one from group 2; under the null the two are interchangeable, so U is near its midpoint.

Because it works on ranks, Mann-Whitney compares whether one distribution is stochastically larger than the other — not whether their means differ. It is the default two-group test for skewed or ordinal data.

Paired Data: Wilcoxon Signed-Rank

When observations come in pairs — before/after, left/right, matched subjects — the Wilcoxon signed-rank test replaces the paired t-test. Compute each pair’s difference, rank the differences by absolute size, then attach the original signs and sum the positive ranks. If the treatment did nothing, positive and negative differences should balance, so the signed-rank sum sits near zero. It uses more information than the even simpler sign test (which counts only the direction of each difference) while still ignoring the raw magnitudes.

Three or More Groups: Kruskal-Wallis and Friedman

For three or more independent groups, the Kruskal-Wallis test is the rank-based one-way ANOVA — it is taught in full in Module 6, Lesson 3, so we only place it in the family here rather than repeating it. Its repeated-measures cousin is the Friedman test: when the same subjects are measured under several conditions (three raters scoring the same wines, say), Friedman ranks within each subject and then tests whether the conditions differ. It is to Kruskal-Wallis what repeated-measures ANOVA is to one-way ANOVA.

Comparing Whole Distributions: Kolmogorov-Smirnov

The tests above compare central tendency. The Kolmogorov-Smirnov (KS) test compares entire distributions by measuring the largest vertical gap between their cumulative distribution functions:

The KS Statistic
D = \max_x \lvert F_1(x) - F_2(x) \rvert
D is the maximum distance between two empirical CDFs (two-sample) or between an empirical CDF and a theoretical one (one-sample, e.g. testing normality). A large D means the distributions differ in shape, spread, or location — not just in their centre.

The General Engine: Permutation Tests

Underneath many of these lies one idea powerful enough to deserve its own name. A permutation test builds the null distribution directly from the data: if the group labels are meaningless under the null, then any shuffling of them is as likely as the observed one. So compute your statistic for the real labels, then for thousands of random relabelings, and see where the real value falls:

Permutation p-value
p = \dfrac{\#\{\text{permutations as extreme as observed}\}}{\#\{\text{all permutations}\}}
The p-value is simply the fraction of shufflings that give a statistic at least as extreme as the observed one. It assumes nothing about the distribution and works for almost any statistic you can compute — the modern, computer-age non-parametric method.
Key Takeaways
  • The study design picks the test: Mann-Whitney U (two independent groups), Wilcoxon signed-rank (paired), Kruskal-Wallis (3+ groups), Friedman (repeated measures).
  • Mann-Whitney compares whether one distribution is stochastically larger — not whether means differ.
  • Kruskal-Wallis is the rank-based one-way ANOVA (taught in Module 6, Lesson 3); Friedman is its repeated-measures version.
  • The Kolmogorov-Smirnov test compares whole distributions via the largest gap between their CDFs — shape and spread, not just centre.
  • Permutation tests build the null by shuffling labels: assumption-free, and applicable to almost any statistic.
Previous Why Non-Parametric? Module Overview Next Non-Parametric Regression