STAT 101
M02 · L01
Module 2

Visualizing Distributions

Before running any statistical test, you need to see your data. Distribution plots reveal shape, spread, and patterns that raw numbers hide.

01 / 13
STAT 101
M02 · L01
The case for charts

Why Visualize?

Two datasets can have identical means and standard deviations yet look completely different. Anscombe’s Quartet famously showed this — always plot your data first.

The golden rule
Plot first, compute second. Numbers summarize; pictures reveal.
02 / 13
STAT 101
M02 · L01
Core tool

Histograms

A histogram groups data into bins and plots the count (or frequency) of observations in each bin. It is the most direct way to see a distribution’s shape.

  • Bins — contiguous, non-overlapping intervals
  • Height — count or relative frequency per bin
  • No gaps — unlike bar charts for categorical data
03 / 13
STAT 101
M02 · L01
The tricky part

Choosing Bin Width

Bin width is the most important histogram decision. Too few bins and you lose detail; too many and noise swamps the signal. Sturges’ rule and Scott’s rule give practical starting points.

Too few bins
Over-smooth
Too many bins
Over-noisy
04 / 13
STAT 101
M02 · L01
Pattern recognition

Distribution Shapes

Histograms reveal the fundamental shapes that data takes in the real world:

  • Symmetric — mirror image around the center
  • Right-skewed — long tail to the right (e.g., incomes)
  • Left-skewed — long tail to the left
  • Bimodal — two peaks (two subgroups in data)
  • Uniform — equal frequency across range
05 / 13
STAT 101
M02 · L01
Shape detail

Skewness

Skewness quantifies the asymmetry of a distribution. A right-skewed (positive) distribution has mean > median > mode. Knowing this tells you which summary statistic to report.

Skewness
\text{Skewness} = \frac{1}{n}\sum_{i=1}^{n}\left(\frac{x_i - \bar{x}}{s}\right)^3
06 / 13
STAT 101
M02 · L01
Smooth curves

Density Plots

Kernel Density Estimation (KDE) smooths the histogram into a continuous curve. The area under the curve always equals 1. KDE removes bin-width dependence and makes shape comparisons easier.

Key parameter
Bandwidth h controls smoothness — analogous to bin width in histograms
07 / 13
STAT 101
M02 · L01
The formula

KDE Formula

KDE places a kernel (usually Gaussian) at each data point, then sums them all up. The result is a smooth estimate of the underlying probability density.

Kernel Density Estimate
\hat{f}(x) = \frac{1}{nh}\sum_{i=1}^{n}K\!\left(\frac{x - x_i}{h}\right)
08 / 13
STAT 101
M02 · L01
Normality check

Q–Q Plots

A Q–Q (quantile-quantile) plot compares your data’s quantiles to those of a theoretical distribution. If the points fall on the diagonal line, your data matches that distribution closely.

Points on line
Normal
Curve above/below
Skewed
09 / 13
STAT 101
M02 · L01
Cumulative view

The ECDF

The Empirical CDF plots the proportion of data points at or below each value. It is a step function from 0 to 1 — no bin-width choice needed, and every data point is represented exactly.

Empirical CDF
F_n(x) = \frac{1}{n}\sum_{i=1}^{n}\mathbf{1}(x_i \le x)
10 / 13
STAT 101
Choosing
Which to use?

Choosing the Right Plot

Each tool has a natural home:

  • Histogram — quick shape overview, large datasets
  • KDE — smooth comparisons across groups
  • Q–Q plot — checking distributional assumptions
  • ECDF — comparing two distributions exactly
11 / 13
STAT 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
STAT 101
Summary
Recap

What you learned

Histograms, density plots, Q–Q plots, and ECDFs each reveal something different about your data. Together they give you a complete picture of any distribution before analysis begins.

13 / 13