Centre Is Only Half the Story
Two towns both report an average annual temperature of 15 °C. In the first, every day of the year sits between 12 and 18 °C. In the second, winter is −20 and summer is +45. The centre is identical and the two places could not be less alike.
What separates them is spread — how far the values scatter around the centre. A summary that reports only a mean is not just incomplete; it invites the reader to imagine a consistency that may not exist. Spread is what tells you how much to trust the centre as a description of any single observation.
This lesson builds five measures of spread and two rules for flagging outliers, all on the same eleven numbers so you can see exactly how each one reacts to the same data.
Every measure of spread answers the same question — how far apart are these numbers? — and each answers it with a different tolerance for extreme values. Choosing one is really a decision about how much weight a single unusual observation should carry.
range · variance · standard deviation · IQR · coefficient of variationOur Running Example
Here are eleven daily commute times in minutes, already sorted:
Ten ordinary mornings and one bad one. That single long commute is not a typo — it is the whole point of the example, and every measure below will be judged by how it reacts to it.
21 23 24 26 28 29 31 33 35 38 53The values sum to 341, so with n = 11 the mean is x̄ = 341 ÷ 11 = 31 minutes exactly. The median is the sixth value, 29 minutes. The mean sits above the median, which by the direction rule from the previous lesson tells you the data is right-skewed — and here that skew comes from one observation.
The Range
The range is the simplest possible measure of spread: the largest value minus the smallest.
For our commutes the range is 53 − 21 = 32 minutes. It is instantly understandable and genuinely useful for a quick sanity check: a range of 32 minutes on a 31-minute average commute already tells you the journey is not reliable.
But look at what the range ignored. Nine of the eleven values played no part in the calculation. Add fifty more perfectly ordinary commutes around 30 minutes and the range does not budge, even though the data is now overwhelmingly consistent. The range is built entirely from the two most extreme observations, which are precisely the two least representative ones.
The range also grows with sample size almost automatically: the more days you record, the better your chance of catching an unusual one. So a range from 11 days and a range from 1,100 days are not comparable quantities. Every measure that follows exists to fix one or both of these problems.
Variance
If we want a measure that uses every observation, the natural move is to average each value's distance from the mean. That instinct is right, but the obvious version of it fails immediately: the deviations from the mean always sum to exactly zero, because that is what being the balance point means. The positives cancel the negatives, every time, for every dataset.
The fix is to square each deviation before averaging. Squaring makes everything positive, and it also makes large deviations count disproportionately — a value twice as far from the mean contributes four times as much. The result is the variance.
Let us compute it on the commutes. The mean is 31, so subtract 31 from every value and square the result.
| xₐ | xₐ − x̄ | (xₐ − x̄)² |
|---|---|---|
| 21 | −10 | 100 |
| 23 | −8 | 64 |
| 24 | −7 | 49 |
| 26 | −5 | 25 |
| 28 | −3 | 9 |
| 29 | −2 | 4 |
| 31 | 0 | 0 |
| 33 | +2 | 4 |
| 35 | +4 | 16 |
| 38 | +7 | 49 |
| 53 | +22 | 484 |
| Total | 0 | 804 |
The middle column sums to zero, exactly as promised. The right column sums to 804, so the variance is s² = 804 ÷ 10 = 80.4.
Notice how lopsided that total is. The single 53-minute commute contributes 484 of the 804 — over 60% of the entire variance from one observation out of eleven. That is squaring doing its work, and it is the reason variance and standard deviation are described as sensitive to outliers.
Standard Deviation
The variance has one awkward property: its units are the square of the data's units. Our commutes are in minutes, so the variance is 80.4 square minutes, which is not a quantity anyone has intuition for.
Take the square root and the units come back. That is the standard deviation, written s for a sample and σ for a population — and it is by far the most widely reported measure of spread in all of statistics.
For the commutes, s = √80.4 = 8.97 minutes. Now the number means something you can act on: a typical morning differs from the 31-minute average by about nine minutes.
This is the exact formula that What is Statistics? put in front of you in Module 1, Lesson 1, under the heading “Sample Standard Deviation.” That lesson stated it, noted that the divisor is n − 1 rather than n, and deferred the reason. This is the lesson where the formula is earned rather than asserted — and the next section is that deferred explanation.
Why n − 1?
Dividing by n − 1 instead of n looks like a fudge, and students reasonably suspect it is one. It is not. Here is what is actually going on, without any proof.
The variance you want is the spread around the population mean μ. But you do not know μ — that is the whole situation described in Population vs. Sample. So you substitute x̄, the mean of the very same data you are now measuring the spread of.
And x̄ is, by construction, the value that makes the sum of squared deviations as small as it can possibly be. Try it: recompute the 804 above using any other centre — 30, or 32 — and the total will be larger. Every time. So measuring your data's spread around its own mean systematically underestimates the spread around the true mean, because your centre has been nudged toward your particular sample.
Dividing by n − 1 rather than n inflates the result by exactly the right amount to cancel that bias out on average. The n − 1 is called the degrees of freedom: eleven observations carry eleven pieces of information, but one of them has already been spent computing x̄, leaving ten free to describe spread. The intuitive check is the extreme case: with a single observation, n − 1 = 0 and the variance is undefined — which is correct, because one data point genuinely tells you nothing whatsoever about variability.
Numerically, on our commutes: 804 ÷ 10 = 80.4 with the correction, and 804 ÷ 11 = 73.1 without it. The uncorrected standard deviation would be 8.55 minutes instead of 8.97 — about 5% too small. With n = 11 the difference is modest; with n = 4 it would be substantial; with n = 1,000 it is negligible. The correction matters most exactly when you have least data.
One consequence to remember: use n − 1 whenever your numbers are a sample, which in practice is almost always. Divide by n only when you genuinely hold every member of the population. Module 5 revisits this as the formal statement that s² is an unbiased estimator of σ².
Quartiles and the IQR
The standard deviation is built on the mean, so it inherits the mean's fragility — we watched one commute supply 60% of it. The alternative is to measure spread by position, exactly as the median measures centre by position.
The median splits sorted data in half. Quartiles split it into quarters. Q1 (the first quartile) is the median of the lower half, Q2 is the median itself, and Q3 (the third quartile) is the median of the upper half. A quarter of the data lies below Q1, and a quarter lies above Q3.
On our eleven commutes, the median is the sixth value, 29. The lower half is 21, 23, 24, 26, 28, whose median is Q1 = 24. The upper half is 31, 33, 35, 38, 53, whose median is Q3 = 35. (With an odd count, the middle value belongs to neither half.)
The interquartile range, or IQR, is simply Q3 − Q1: the width of the band containing the middle 50% of the data.
Here IQR = 35 − 24 = 11 minutes. Now the crucial property: change the 53 to 200 and the IQR does not move at all. Q1 and Q3 are positions, and moving the largest value further out does not change which value is fourth from the top. The IQR is a robust measure of spread in the same way the median is a robust measure of centre — and that is why it is the foundation of the outlier rule later in this lesson.
The Five-Number Summary
Put the extremes together with the three quartiles and you have the five-number summary: minimum, Q1, median, Q3, maximum. Five numbers that describe centre, spread and shape simultaneously, and that survive outliers in the middle three.
| Statistic | Value | Reads As |
|---|---|---|
| Minimum | 21 | The best morning of the eleven |
| Q1 | 24 | A quarter of mornings were faster than this |
| Median | 29 | A typical morning |
| Q3 | 35 | A quarter of mornings were slower than this |
| Maximum | 53 | The worst morning of the eleven |
Read the gaps and the shape appears without a single chart. From the minimum to Q1 is 3 minutes; from Q1 to the median is 5; from the median to Q3 is 6; from Q3 to the maximum is 18. The spacing widens sharply as you move right, which is what right-skew looks like in numbers.
Building a Box Plot, Step by Step
A box plot is the five-number summary drawn to scale. Other lessons in this course refer to box plots as a tool you already have — Effective Data Communication in Module 2 recommends one for comparing distributions across groups, and Confidence Intervals for a Mean in Module 5 tells you to check one before trusting a t-interval on a small sample. Here is how you actually construct it.
Step 1 — sort the data. Everything that follows is positional, so nothing works until the values are in order. Ours are already sorted: 21 through 53.
Step 2 — write down the five-number summary. Minimum 21, Q1 24, median 29, Q3 35, maximum 53.
Step 3 — compute the IQR and the two fences. IQR = 11, so 1.5 × IQR = 16.5. The lower fence is Q1 − 16.5 = 24 − 16.5 = 7.5, and the upper fence is Q3 + 16.5 = 35 + 16.5 = 51.5. The fences are not drawn on the finished plot; they exist to decide where the whiskers stop.
Step 4 — draw the box. A rectangle spanning Q1 to Q3, from 24 to 35. The width of that box is the IQR, which is why a box plot shows you a robust measure of spread at a glance.
Step 5 — draw the median line. A line across the box at 29. It sits left of centre in our box (5 minutes from Q1, 6 from Q3), and an off-centre median line is your visual signal of skew.
Step 6 — draw the whiskers. Each whisker extends from the box to the most extreme value that is still inside the fence — not to the fence itself. Below, nothing is under 7.5, so the lower whisker reaches the minimum, 21. Above, 53 exceeds the 51.5 fence, so the upper whisker stops at the next value down, 38.
Step 7 — plot the stragglers individually. Any value outside a fence is drawn as its own point. Our 53 becomes a single dot beyond the end of the upper whisker: visible, labelled, and not silently averaged into anything.
The finished plot reads left to right as 21 — [24 | 29 | 35] — 38 •53. Four facts in one glance: the middle half spans 24 to 35, a typical morning is 29, ordinary mornings top out around 38, and one morning was genuinely unusual. The companion slide deck has this box plot on a canvas with a slider on the worst commute — drag it and watch the point cross the fence.
The Coefficient of Variation
Every measure so far carries the units of the data, which makes comparing spreads across different quantities meaningless. Is a standard deviation of 8.97 large? You cannot possibly say without knowing that the average is 31 — and you certainly cannot compare it to a standard deviation measured in kilograms.
The coefficient of variation solves this by dividing the standard deviation by the mean. The units cancel, and what is left is spread expressed as a proportion of the average, usually reported as a percentage.
For our commutes, CV = 8.97 ÷ 31 = 0.289, or 28.9%. The typical morning deviates from the average by nearly 30% of the average. That single number is interpretable with no knowledge of minutes at all.
Here is the cleanest demonstration that it is genuinely unitless. Adult heights might have a mean of 170 cm and a standard deviation of 7 cm. Record the same people in metres and the standard deviation becomes 0.07 — a hundred times smaller, describing identical people. But the CV is 7 ÷ 170 = 4.1% in centimetres and 0.07 ÷ 1.70 = 4.1% in metres. Change the unit and the standard deviation changes; the coefficient of variation does not.
That is what makes cross-scale comparison possible. Adult body temperature has a mean near 36.8 °C and a standard deviation around 0.4 °C, giving a CV of about 1.1%. Our commute's CV is 28.9%. Relative to its own typical size, the commute is roughly twenty-six times more variable than body temperature — a statement you could not make from the raw standard deviations of 0.4 and 8.97, which live in incomparable units.
Two cautions. The CV requires a meaningful, non-zero mean, so it is unsuitable for anything centred near zero (the denominator explodes) and for interval scales where zero is arbitrary — the CV of a temperature in Celsius and the CV of the same temperature in Fahrenheit disagree, because those scales have different zeros. And because it is built from x̄ and s, it inherits their sensitivity to outliers.
Outliers: the 1.5 × IQR Fence
An outlier is an observation that sits far enough from the rest of the data to deserve a second look. Notice the wording: deserving a look, not deserving deletion. An outlier can be a measurement error, a data-entry slip, or the single most important observation in the study.
The standard univariate rule is the one we already used to draw the whiskers. Build a fence 1.5 IQRs beyond each quartile, and flag anything outside it.
For the commutes: 7.5 and 51.5. Nothing falls below 7.5. The 53-minute morning exceeds 51.5, so it is flagged — the same conclusion the box plot drew geometrically, because it is literally the same calculation.
Why 1.5? It is a convention, chosen by John Tukey when he invented the box plot, and its justification is empirical rather than deep: on data that really is roughly normal, these fences flag only about 0.7% of observations, so a flag is unusual enough to be worth investigating without drowning you in false alarms. Some analysts use 3.0 × IQR to mark “far outliers” separately.
The rule's great virtue is robustness. It is built from Q1, Q3 and the IQR, none of which move when you push the extreme value further out. A cluster of outliers cannot conspire to hide itself by widening the very fence meant to catch it — which, as the next section shows, is exactly what happens with the alternative.
Outliers: the z-Score Rule
The second rule counts standard deviations. The z-score of an observation is its distance from the mean expressed in units of s.
The usual convention flags |z| > 3, sometimes |z| > 2 on small samples. For our 53-minute commute: z = (53 − 31) ÷ 8.97 = 2.45. Under the |z| > 3 rule that morning is not an outlier — while the IQR fence flagged it clearly.
The two rules disagree, and understanding why is the most useful thing in this section. Both x̄ and s are computed including the suspect value. The 53 dragged the mean up from 28.8 to 31, shortening its own distance from the centre, and it supplied 60% of the sum of squares, inflating the very s it is being divided by. It sabotaged the test applied to it.
Quantify that. Set the 53 aside and the remaining ten commutes have a mean of 28.8 and a standard deviation of 5.49. Measured against the data it does not belong to, the 53-minute morning scores z = (53 − 28.8) ÷ 5.49 = 4.41 — comfortably past any threshold. The same observation is a 2.45 or a 4.41 depending only on whether it was allowed to vote on its own trial. Statisticians call this effect masking.
Masking gets worse, not better, as outliers pile up: the more extreme values there are, the more they inflate s, and the harder it becomes for any of them to exceed a fixed z threshold. Robust variants exist — replace x̄ with the median and s with the median absolute deviation — and are worth reaching for. But the plain z-score rule is what most people mean by “the z-score rule”, so know its weakness.
| Rule | Built From | Verdict | Robust? |
|---|---|---|---|
| 1.5 × IQR fence | Q1, Q3, IQR | 53 > 51.5 — flagged | Yes: the fence does not move |
| z-score, |z| > 3 | x̄ and s | z = 2.45 — not flagged | No: the value inflates its own s |
| z-score, value excluded | x̄ and s of the other 10 | z = 4.41 — flagged | Better, but needs you to guess first |
Both rules here are univariate: they judge one variable at a time. Once a point is being judged against a fitted model rather than against a single column, the machinery changes entirely — standardized and studentized residuals, leverage, and Cook's distance. Regression Diagnostics in Module 7 develops all of that, and there is no need to duplicate it here; the rules above are what you use when all you have is a column of numbers.
Finally, the discipline that matters more than either rule: never delete an outlier for being inconvenient. Investigate it. If it is a recording error, fix or remove it and say so. If it is real, keep it, and report your analysis both with and without it so your reader can see how much the conclusion depends on one point. A 53-minute commute may be the morning that explains everything.
- Two datasets with the same mean can be completely different: spread tells you how much to trust the centre.
- The range (53 − 21 = 32) uses only two observations and grows with sample size, so it is a sanity check rather than a measure.
- Deviations from the mean always sum to zero, which is why variance squares them first: s² = 804 ÷ 10 = 80.4, and s = 8.97 minutes.
- n − 1 corrects for the fact that x̄ minimises the sum of squared deviations, so measuring spread around it underestimates the truth — 804 ÷ 11 would give 8.55, about 5% too small.
- Q1 = 24 and Q3 = 35 give IQR = 11, and the IQR does not move if you change the largest value to 200. That robustness is its whole purpose.
- The five-number summary 21 / 24 / 29 / 35 / 53 becomes a box plot: box from Q1 to Q3, line at the median, whiskers to the last value inside the fences, stragglers as their own points.
- CV = s ÷ x̄ = 28.9% is unitless, so it survives a change from centimetres to metres and lets you compare a commute's variability with body temperature's 1.1%.
- The 1.5 × IQR fence flags our 53 while the z-score rule does not (z = 2.45), because the outlier inflates the very s it is judged against — exclude it and z jumps to 4.41. That is masking.