Reading
Stories Mode

Visualizing Distributions

~12 min read Lesson 1 of 3 in Module 2

Always Plot First

Before running any statistical test or computing any summary statistic, there is one rule every data analyst should follow: plot your data first. Visualizing a distribution tells you things that numbers alone cannot — the shape, the presence of outliers, whether subgroups exist, and whether your data meets the assumptions required by the analysis you plan to run.

The statistician Francis Anscombe made this point memorably in 1973 with what is now called Anscombe’s Quartet: four datasets with nearly identical means, variances, and correlations, yet wildly different visual shapes. The lesson has not changed: summary statistics can be deceiving; pictures reveal the truth.

Anscombe’s Quartet

Four datasets each have the same mean (~7.5), variance (~4.12), correlation (~0.816), and regression line. Yet one is linear, one is curved, one has an outlier, and one has almost no variation in X. Without plotting, you would treat them identically.

Plot first, compute second. Numbers summarize; pictures reveal.

Histograms

The histogram is the workhorse of distribution visualization. It divides the range of a variable into contiguous, non-overlapping intervals called bins, counts how many observations fall in each bin, and draws a bar whose height reflects that count. Unlike bar charts for categorical data, histogram bars touch each other — the continuity signals that the underlying variable is continuous.

Reading a histogram is intuitive: tall bars indicate where values cluster; short bars indicate sparse regions; gaps suggest possible subgroups or measurement artifacts.

Choosing Bin Width

Bin width is the most consequential choice in building a histogram. Too few bins and you over-smooth the data, hiding important features. Too many bins and random noise dominates, obscuring genuine patterns.

Sturges’ rule suggests k = 1 + log₂(n) bins as a starting point. Scott’s rule sets bin width to 3.49σn−1/3. Both are heuristics — always try a few different widths.

Distribution Shapes

One of the most valuable things a histogram can reveal is the shape of a distribution. Shape is not merely aesthetic — it determines which statistical methods are appropriate and which summary statistics are most informative.

Symmetric distributions have a mirror-image appearance around their center. The normal distribution is the canonical example. For symmetric data, the mean, median, and mode are close together, and either the mean or median is a good summary.

Right-skewed (positive) distributions have a long tail stretching to the right. Income, wealth, and city populations often follow this pattern: most values cluster at the low end, but a few extreme values pull the tail far to the right. Here the mean is pulled above the median — the median is a more robust summary.

Left-skewed (negative) distributions mirror this, with a long tail to the left. Test scores in an easy exam (most students do well, a few score very low) commonly exhibit left skew.

Bimodal distributions show two distinct peaks, signaling that the data may contain two subpopulations. Combining morning and afternoon commute times, or heights from a mixed-gender sample without separation, often produces bimodality.

Uniform distributions show roughly equal frequency across all values. A fair die rolled many times should produce a nearly uniform histogram.

Quantifying Skewness

Visual inspection tells us the direction of skew, but we can quantify it with the skewness statistic — the standardized third central moment. A value near zero indicates symmetry; positive values indicate right skew; negative values indicate left skew.

Skewness
\text{Skewness} = \frac{1}{n}\sum_{i=1}^{n}\left(\frac{x_i - \bar{x}}{s}\right)^3
The third standardized moment: positive values indicate a right tail, negative values a left tail, and zero indicates symmetry

Density Plots and KDE

A density plot replaces the histogram with a smooth continuous curve. It is produced by Kernel Density Estimation (KDE): for each data point, a kernel function (usually a Gaussian) is placed on that point, and all kernels are summed. The result is a smooth estimate of the underlying probability density function.

The key advantage of KDE over histograms is that it removes the arbitrary bin-boundary problem. Moving a boundary slightly in a histogram can change which bar a value falls into; KDE has no such discontinuity. The area under a KDE curve always equals 1, so it can be interpreted as a probability density.

Kernel Density Estimate
\hat{f}(x) = \frac{1}{nh}\sum_{i=1}^{n}K\!\left(\frac{x - x_i}{h}\right)
Each data point xᵢ contributes a kernel K centered on it; bandwidth h controls smoothness (large h = over-smooth, small h = under-smooth)

The bandwidth h plays the same role as bin width in a histogram. Silverman’s rule of thumb — h = 1.06σn−1/5 — works well when the data are approximately normal. For multi-modal data, smaller bandwidths often reveal more structure.

Q–Q Plots: Checking Normality

Many statistical methods assume that data follow a normal distribution. A quantile-quantile (Q–Q) plot provides a visual test of this assumption. It plots the quantiles of your data on the vertical axis against the corresponding quantiles of a normal distribution on the horizontal axis.

If the data are perfectly normal, all points fall exactly on a diagonal line. Deviations from the line reveal departures from normality:

Points curve above the line in the right tail: the data has heavier tails than normal (leptokurtic, like a t-distribution).

Points curve below the line in the right tail: the data has lighter tails than normal (platykurtic).

S-shaped curves: indicate skewness. An upward S indicates right skew; a downward S indicates left skew.

Reading a Q–Q Plot

Points on the line: Data match the reference distribution (e.g., normal).

Points curve outward at both ends: Heavy tails — more extreme values than expected.

Points curve inward at both ends: Light tails — fewer extreme values than expected.

The Empirical CDF

The Empirical Cumulative Distribution Function (ECDF) is perhaps the most honest visualization of a distribution: it makes no binning or smoothing decisions, and every data point is represented exactly. The ECDF at any value x gives the proportion of data points that are less than or equal to x.

Empirical CDF
F_n(x) = \frac{1}{n}\sum_{i=1}^{n}\mathbf{1}(x_i \le x)
The indicator function 1(xᵢ ≤ x) equals 1 when observation xᵢ is at most x, and 0 otherwise — the ECDF counts the fraction of data at or below each value

The ECDF is a step function: it starts at 0, jumps by 1/n at each data point, and reaches 1 at the largest observation. ECDFs are particularly useful for comparing two distributions: plot both on the same axes and the gap between them at any x-value shows the difference in cumulative probability at that point.

Stem-and-Leaf Plots

Before computer graphics, statisticians used stem-and-leaf plots to display distributions while preserving individual data values. Each observation is split into a stem (leading digit(s)) and a leaf (trailing digit). Rotating a stem-and-leaf plot 90° produces something that looks like a histogram — but every exact value is retained.

Stem-and-leaf plots work best for small datasets (under ~100 observations) and are rarely seen in modern software, but they remain a useful mental model: they remind us that a histogram is an aggregation of individual data points, not a smooth mathematical function.

Choosing the Right Tool

Each visualization has a natural home:

Histogram: Best for a quick first look at a single variable. Use when the dataset is large enough that individual points would clutter a plot.

KDE: Best for smooth comparisons — especially when overlaying multiple groups on the same plot. The smooth curve is easier to compare visually than overlapping histograms.

Q–Q plot: Best for checking whether data meet a specific distributional assumption (usually normality) before applying a parametric test.

ECDF: Best when you need an exact comparison between two distributions, or when you want to read off percentiles precisely without smoothing bias.

In the next lesson, we will move from single-variable distributions to visualizing relationships between two variables — scatter plots, correlation, and heatmaps — the tools that reveal how variables interact.

Key Takeaways
  • Always plot your data before computing statistics — Anscombe’s Quartet shows that identical summaries can hide completely different shapes.
  • Histograms reveal distribution shape but depend on bin-width choice; try multiple widths before settling on one.
  • KDE removes bin-boundary artifacts and is ideal for smooth group comparisons; bandwidth plays the same role as bin width.
  • Q–Q plots test whether data match a theoretical distribution; deviations from the diagonal line reveal skew and tail behavior.
  • The ECDF is the most exact representation — no smoothing, every point visible — and is ideal for precise distribution comparisons.
Previous Measures of Spread Module Overview Next Lesson Visualizing Relationships