Reading
Stories Mode

Measures of Central Tendency

~14 min read Lesson 4 of 5 in Module 1

One Number to Stand for Many

You have 500 commute times, 10,000 salaries, or a year of daily temperatures. Nobody reads 500 numbers. The first question anyone asks is the same: where is the middle? A measure of central tendency answers it — one value that stands in for the whole batch.

There is more than one honest answer to “where is the middle,” and that is the whole point of this lesson. The arithmetic mean is the default everywhere, but it is the wrong default for growth rates, for speeds, for skewed data, and for categories. Choose the wrong measure and your summary does not throw an error — it quietly lies.

By the end of this lesson you will have six tools instead of one, and a rule for picking between them that depends on what the numbers actually are: sums, ratios, rates, ranks, or names.

The Core Idea

A measure of central tendency compresses a whole dataset into a single number. Which measure is correct is not a matter of taste — it is determined by the kind of quantity you are averaging.

mean · weighted mean · geometric mean · harmonic mean · median · mode

The Arithmetic Mean

The arithmetic mean is the one everybody already knows: add all the values, divide by how many there are. It is written x̄ (“x-bar”) for a sample and μ for a population — the notation you met in Population vs. Sample.

Arithmetic Mean
\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i
Sum every observation, divide by the count n — the sample mean x̄

Take five quiz scores: 72, 85, 90, 68, 95. They sum to 410, and 410 ÷ 5 = 82. That is the arithmetic mean, and for a small set of well-behaved scores it is exactly the right summary.

The mean has a physical meaning worth carrying around. Lay the values out on a number line as equal weights: the mean is the point where a fulcrum would balance them. Every observation pushes on it, and each one pushes in proportion to how far it sits from the balance point.

That is also the mean's weakness. One value placed far enough out tips the whole beam, and no amount of good data on the other side fully resists it. Hold on to that image — it is exactly what the median section is about.

The Weighted Mean

The plain mean treats every value as equally important. Very often they are not, and averaging as though they were produces a number that is simply wrong.

Here is the case to remember. A course has two sections. Section A has 10 students and averages 90. Section B has 40 students and averages 70. What is the course average?

Averaging the two averages gives (90 + 70) ÷ 2 = 80. But there are 50 students in the course, and 40 of them are in the 70 group. The 90 belongs to a fifth of the class and the naive average hands it half the vote.

Weighted Mean
\bar{x}_w = \frac{\sum_{i=1}^{n} w_i x_i}{\sum_{i=1}^{n} w_i}
Each value carries a weight wₐ — the plain mean is the special case where every weight is equal

Weight each section by its size and you get (10 × 90 + 40 × 70) ÷ 50 = 3700 ÷ 50 = 74. Six points lower than the naive answer. That is not a rounding difference; it is the gap between a cohort you would call solid and one you would call struggling.

The rule is unconditional: whenever you average summaries of groups of unequal size, weight by group size. The unweighted average of averages is one of the most common arithmetic errors in published statistics, and it is invisible once the group sizes are dropped from the report.

Weights do not have to be counts. A course grade might be 20% homework, 30% midterm, 50% final. If you score 95, 78 and 84, the plain mean is 85.7 but your grade is 0.20 × 95 + 0.30 × 78 + 0.50 × 84 = 84.4. When the weights already sum to 1, the denominator disappears and the weighted mean is just the dot product of scores and weights.

The Geometric Mean

Some numbers do not add — they multiply. Growth rates, interest rates, inflation, index numbers, ratios of any kind: chain two periods together and you multiply their factors, you do not sum them. Averaging them arithmetically gives an answer that contradicts the data.

Watch it fail. An investment gains 50% in year one and loses 40% in year two. The arithmetic mean of +50% and −40% is +5% per year, which says you made money.

You did not. Convert the returns to growth factors: 1.50 and 0.60. Chain them: 1.50 × 0.60 = 0.90. Two years in, you hold 90 cents for every dollar you started with. You are down 10%, and the arithmetic mean reported a gain.

The geometric mean is the fix. Multiply the values and take the n-th root.

Geometric Mean
G = \left(\prod_{i=1}^{n} x_i\right)^{1/n} = \sqrt[n]{x_1 x_2 \cdots x_n}
The n-th root of the product — equivalently, the arithmetic mean of the logarithms, exponentiated back

For our two factors: √(1.50 × 0.60) = √0.90 = 0.9487, which is −5.13% per year. Check it the only way that matters: 0.9487² = 0.90, exactly the two-year outcome. The geometric mean is the constant rate that reproduces the actual end result — that is its definition in one sentence, and it is why finance quotes compound annual growth rates this way.

Two properties are worth memorising. First, the geometric mean is never larger than the arithmetic mean, and the two are equal only when every value is identical. Second, it needs strictly positive values — a single zero drags the product to zero, and a negative value makes the root meaningless. Average growth factors (1.50), not growth percentages (+50%), and this is never a problem.

You will meet this term again in Module 7. When a regression is fitted to log(Y) and the prediction is transformed back, what comes out is the geometric mean of Y, not the arithmetic mean — a remark in Regression Diagnostics that only makes sense once you have the definition above.

The Harmonic Mean

A third family of quantities breaks the arithmetic mean in yet another way: rates. Speed is kilometres per hour, fuel economy is litres per 100 km, throughput is bits per second. These are ratios with a denominator, and averaging them as plain numbers ignores the denominator entirely.

The classic case: you drive 60 km to a site at 30 km/h and the same 60 km home at 60 km/h. What was your average speed?

The arithmetic mean says 45 km/h. Now compute it honestly. The outbound leg took 60 ÷ 30 = 2 hours. The return took 60 ÷ 60 = 1 hour. You covered 120 km in 3 hours, so your average speed was 40 km/h — five short of the arithmetic answer.

The harmonic mean gets it right in one step: n divided by the sum of the reciprocals.

Harmonic Mean
H = \frac{n}{\sum_{i=1}^{n} \dfrac{1}{x_i}}
The reciprocal of the mean of the reciprocals — the correct average for rates over a shared denominator

Substituting: 2 ÷ (1/30 + 1/60) = 2 ÷ (3/60) = 2 ÷ 0.05 = 40 km/h. Exactly the honest answer, with no need to invent a distance.

The reason it works is worth stating plainly: you spend more time at the slow speed than at the fast one, so the slow leg deserves more weight in the average. The harmonic mean applies that weighting automatically. The same logic makes it the right average for price-to-earnings ratios across a portfolio and for the F₁ score in machine learning, which is the harmonic mean of precision and recall precisely so that one very low component cannot be hidden by one very high one.

The Three Means Compared

For any set of positive numbers that are not all identical, the three means always fall in the same order: harmonic < geometric < arithmetic. Our two speeds demonstrate it.

The Three Means on 30 and 60
Mean Computed As Result Use When
Harmonic 2 ÷ (1/30 + 1/60) 40.00 Rates and ratios over a shared denominator
Geometric √(30 × 60) 42.43 Growth factors, index numbers, multiplicative change
Arithmetic (30 + 60) ÷ 2 45.00 Additive quantities: counts, lengths, scores, money

The ordering is also a free sanity check. If you compute a geometric mean and it lands above the arithmetic mean of the same numbers, you have made an arithmetic slip — that outcome is impossible.

The Median

The median abandons arithmetic entirely and asks about position: sort the values, then take the middle one. With an odd count there is a single middle value. With an even count there are two, and the median is their average.

Median
\text{median} = \begin{cases} x_{\left(\frac{n+1}{2}\right)} & n \text{ odd} \\[6pt] \dfrac{x_{\left(\frac{n}{2}\right)} + x_{\left(\frac{n}{2}+1\right)}}{2} & n \text{ even} \end{cases}
x⁽ₖ⁾ denotes the k-th value after sorting — the median depends only on order, never on magnitude

Five houses on a street sell for 240, 255, 270, 290 and 310 thousand. The median is the third value, 270. The mean is 1365 ÷ 5 = 273. They agree closely, because the data is well behaved.

Now a mansion at the end of the street sells for 2,000 thousand instead of 310. The mean becomes 3055 ÷ 5 = 611 — a figure that describes no house on the street and would mislead every buyer who read it. The median is still 270. It did not move at all.

Statisticians measure this quality with a breakdown point: the fraction of the data an adversary must corrupt before your summary can be pushed anywhere they like. The median's breakdown point is 50% — you can wreck almost half the values and it still reports something sensible. The mean's breakdown point is 0%. One observation, moved far enough, takes the mean with it.

With an even count, take the two middle values and average them: for 240, 255, 270 and 290 the median is (255 + 270) ÷ 2 = 262.5. Note that the median need not be a value that actually occurs in the data.

Because the median needs only order, it works on ordinal data, where the mean is meaningless. Recall the satisfaction scale from Types of Data: the average of “satisfied” and “very satisfied” is not a thing, but the middle response of 400 customers absolutely is.

The Mode

The mode is the value that occurs most often. It requires no arithmetic and no ordering — only counting — which makes it the only measure of central tendency that survives on nominal data.

Ask 100 customers how they prefer to be contacted: 46 say email, 31 say phone, 23 say SMS. The mode is email. You cannot compute a mean here, because email and phone are names rather than numbers. You cannot compute a median either, because nominal categories have no order to have a middle of — as Types of Data established, that is what makes them nominal.

So on nominal data the mode is not a weaker fallback. It is the answer, and reaching for anything else is a category error.

On continuous data the mode gets slippery, because typically no value repeats at all — every measured temperature is unique to the last decimal. There, “the mode” means the peak of the distribution, which in practice means the tallest bar of a histogram. That makes it dependent on your bin width, a sensitivity that Visualizing Distributions in Module 2 takes up in detail.

The mode has one talent the mean and median lack: it can report more than one centre. A bimodal distribution has two peaks, and two peaks usually mean two populations have been mixed together — adult heights pooled across both sexes, or response times from cached and uncached requests. In that situation any single centre is a fiction, and the honest summary is that there is not one centre. Only the mode can say so.

Mean vs. Median Under Skew

This is the section that makes the choice matter, and it has a definite direction you can rely on. When a distribution is skewed, its tail stretches further on one side than the other, and the mean is pulled toward the long tail while the median stays put.

Right-skewed (positively skewed) data has a long right tail, so the mean sits above the median. Income, wealth, city populations, insurance claims, waiting times and file sizes all behave this way: most values bunch at the low end and a few very large ones stretch out to the right.

Left-skewed (negatively skewed) data mirrors it, with a long left tail, so the mean sits below the median. Scores on an exam most students found easy look like this, as does age at death in a country with low infant mortality.

For a symmetric distribution the two coincide, at least to within sampling noise. So the sign of x̄ − median is a free, one-line skew diagnostic: positive means right-skewed, negative means left-skewed, near zero means roughly symmetric.

The Diagnostic on Our Street

The five house prices with the mansion included gave a mean of 611 and a median of 270. Their difference is +341 — strongly positive, so strongly right-skewed, and here the entire skew is produced by a single observation.

x̄ > median → right-skewed  ·  x̄ < median → left-skewed

The practical consequence is to report both. Average household income and median household income differ substantially in every country on earth, and which one a headline chooses is a decision about what story to tell. Seeing both at once tells you not only where the centre is but how lopsided the distribution around it must be.

The companion slide deck for this lesson has a histogram with a skew slider. Drag it and watch the mean and median separate and swap places — the relationship above is much easier to trust once you have moved it yourself.

Choosing a Measure

Putting it together, the choice is driven by the kind of data you hold, not by convention or convenience.

Which Measure, and Why
Your Data Reach For Because
Nominal categories Mode No order and no arithmetic are available
Ordinal ranks Median (mode also fine) Order is meaningful but the gaps are not equal
Symmetric numeric Mean Uses every observation, and is the most precise estimator
Skewed or outlier-prone Median The mean is dragged toward the long tail
Growth rates and ratios Geometric mean The quantities multiply rather than add
Rates over a shared denominator Harmonic mean Weights each rate by the time or quantity it applies to
Groups of unequal size Weighted mean Each group should count in proportion to its size

Knowing the centre is only half of a summary. Two datasets can share an identical mean and median and still describe completely different worlds — one tightly clustered, one wildly scattered. The next lesson builds the other half: range, variance, standard deviation, the IQR, the five-number summary, the box plot, and rules for deciding when a value is genuinely an outlier.

Key Takeaways
  • The arithmetic mean is the balance point of the data: it uses every observation, which is both its strength and its fragility.
  • Averaging summaries of unequal-sized groups requires a weighted mean — two sections of 10 and 40 students averaging 90 and 70 give 74, not 80.
  • Use the geometric mean for anything that multiplies: +50% then −40% is a 10% loss over two years, not a 5% annual gain.
  • Use the harmonic mean for rates over a shared denominator: 60 km at 30 km/h and 60 km at 60 km/h averages 40 km/h, not 45.
  • For positive values that are not all equal, harmonic < geometric < arithmetic, always — a free check on your arithmetic.
  • The median has a 50% breakdown point and the mean has 0%: one mansion moved a street's mean from 273 to 611 and left the median untouched.
  • The mode is the only central measure available for nominal data, and the only one that can report two centres when two populations have been mixed.
  • Under skew the mean is pulled toward the long tail: x̄ > median means right-skewed, x̄ < median means left-skewed. Report both.
Previous Population vs. Sample Module Overview Next Lesson Measures of Spread