From One Variable to Two
The previous lesson showed how to visualize a single variable — its shape, spread, and tail behavior. But data analysis rarely ends there. Most interesting questions involve relationships: does studying more lead to higher grades? Does temperature predict ice cream sales? Does advertising spend drive revenue? These questions demand charts that show two (or more) variables simultaneously.
Relationship visualization serves two purposes: exploration (finding patterns we did not anticipate) and communication (showing patterns we already know exist to others). The tools covered in this lesson — scatter plots, line plots, bubble charts, heatmaps, and pair plots — cover the full range from two variables to many.
Scatter Plots
The scatter plot is the fundamental tool for visualizing the relationship between two quantitative variables. Each observation becomes a point; its horizontal position encodes one variable and its vertical position encodes the other. The result is a cloud of points whose shape tells you everything about the relationship.
When reading a scatter plot, you are looking for four things:
Direction: Do the points tend to rise from left to right (positive association) or fall (negative association)? Or is there no clear direction?
Form: Is the relationship linear (points scattered around a straight line) or nonlinear (curved, U-shaped, or more complex)?
Strength: How tightly do the points cluster around the trend? A tight cluster means a strong relationship; a widely dispersed cloud means a weak one.
Unusual features: Are there outliers — points far from the main cloud? Are there clusters or subgroups within the data?
A single outlier can dramatically change the apparent correlation in a scatter plot. Before reporting any relationship, look for influential points that are far from the main cloud in both x and y. These may represent data errors, measurement artifacts, or genuinely exceptional observations that deserve separate analysis.
Correlation: Numerical Summary of a Scatter Plot
The scatter plot shows the visual picture; the correlation coefficient provides a number that summarizes the direction and strength of a linear relationship. The most widely used measure is Pearson’s r, which ranges from −1 (perfect negative linear relationship) to +1 (perfect positive linear relationship), with 0 indicating no linear relationship.
Pearson’s r has an intuitive interpretation: it is the average product of the standardized x and y values. When both variables tend to be above their means together (and below together), the products are positive and r is positive. When one tends to be above while the other is below, the products are negative and r is negative.
Rough guidelines for interpreting r: |r| < 0.3 is weak, 0.3 – 0.6 is moderate, and |r| > 0.6 is strong. But these are domain-dependent — in physics, r = 0.9 might be expected; in psychology or social science, r = 0.4 might be considered strong.
This is perhaps the most important warning in all of statistics. Ice cream sales and drowning rates are positively correlated — not because ice cream causes drowning, but because both rise in summer. A high correlation tells you two variables move together; it says nothing about why. Establishing causation requires experimental design, temporal ordering, and ruling out confounders.
Correlation measures association, not causation. Always ask: could there be a third variable driving both?Spearman Rank Correlation
Pearson’s r captures linear relationships. When the relationship is monotone but not linear — or when the data contain outliers or violate normality assumptions — Spearman’s rank correlation (ρ) is more appropriate. It simply applies Pearson’s formula to the ranks of the data rather than the raw values.
Spearman’s ρ is +1 if the rankings are in perfect agreement, −1 if they are in perfect reversal, and 0 if the rankings are unrelated. Since it uses ranks, it is robust to outliers and works for ordinal data as well as continuous measurements.
Line Plots: Trends Over Time
When one of the two variables is time, the scatter plot gives way to the line plot. Connecting the data points with a line makes temporal ordering explicit and helps the eye follow trends, cycles, and sudden changes. The x-axis always represents time (date, year, quarter, hour, etc.) and the y-axis the quantity of interest.
Line plots excel at revealing:
Trends: A sustained upward or downward movement over time — GDP growing, temperature rising, costs falling.
Seasonality: Regular, repeating patterns tied to time of year, day of week, or hour of day. Retail sales spike every December; electricity demand peaks every weekday morning.
Step changes and anomalies: Sudden jumps that mark policy changes, system failures, or external shocks. A pandemic, a product launch, or a regulation often appears as a sharp kink in a line plot.
When multiple lines are plotted on the same axes, the chart shows how the gap between series changes over time. Keep the number of lines to five or fewer; beyond that, the chart becomes unreadable. If you have many series, consider small multiples (one chart per series, sharing the same axes).
Bubble Charts: Three Variables in Two Dimensions
A bubble chart extends the scatter plot by encoding a third quantitative variable as the size of each point. The x and y axes show two variables; the area of the circle (bubble) encodes the third. Hans Rosling’s famous animated bubble charts of global development — with countries as bubbles, x = income, y = life expectancy, and size = population — showed the world that rich, healthy, and populous are not the same thing.
Bubble charts work best when:
• The third variable has a wide range and meaningful differences in magnitude matter.
• The number of observations is moderate (5–50). Too many bubbles create overlap that hides individual values.
• Bubble size encodes area, not radius. The human eye perceives area, so map the variable to πr², not r directly.
A common mistake is mapping a variable to the bubble’s radius rather than its area. If country A has twice the population of country B, its bubble should have twice the area, not twice the radius. A radius-doubled bubble has four times the area — a misleading exaggeration. Always verify that your charting tool maps to area by default, or compute radius = √(value / π).
Heatmaps: Many Variables at Once
When the number of variables grows beyond three, two-dimensional plots begin to struggle. A heatmap encodes a matrix of values using color intensity — typically a gradient from cool (low values) to warm (high values), or from one color to another for diverging data. Each cell of the matrix gets a color, and the reader’s eye scans for patterns.
The most common use of heatmaps in statistics is the correlation matrix: a square grid where cell (i, j) shows the correlation between variable i and variable j. The diagonal is always 1.0 (each variable perfectly correlates with itself). Off-diagonal cells reveal which pairs of variables are strongly related — and in what direction.
Heatmaps are also used for:
Time × category matrices: Sales by product by month. One axis is time (month), the other is category (product), and the color encodes the value.
Geographic data: Population density, temperature, or pollution level mapped onto a geographical grid.
Confusion matrices: In machine learning, heatmaps visualize how often each predicted class matches each true class.
Pair Plots: All Pairwise Relationships
When exploring a dataset with many variables, you want to see all pairwise scatter plots at once. A pair plot (also called a scatter plot matrix or SPLOM) creates a grid where each row and column represents one variable. The cell at row i, column j shows a scatter plot of variable i on the y-axis vs. variable j on the x-axis. The diagonal typically shows either a histogram or a KDE for each variable by itself.
Pair plots are an invaluable first-pass tool in exploratory data analysis. By glancing at a single figure, you can spot which variable pairs show strong linear relationships (candidates for regression), which pairs show no relationship (variables that may be independent), and which pairs show nonlinear or clustered patterns (signals that the data may contain subgroups or require transformation).
The main limitation is scale. With p variables, a pair plot has p² cells. For p = 5, that is 25 cells — manageable. For p = 20, that is 400 cells — unreadable. When you have more than 8–10 variables, use the correlation heatmap instead, or pre-select the most interesting variables based on domain knowledge before building the pair plot.
Choosing the Right Chart
Scatter plot: Two continuous variables, no time ordering, moderate to large sample size. The default choice for relationship exploration.
Line plot: One axis is time. Use when the temporal sequence matters and you want to show trends, cycles, or sudden changes.
Bubble chart: Three continuous variables, small to moderate number of observations. Good when the third variable’s magnitude has a meaningful interpretation.
Heatmap: Many variables (correlation matrix) or a value defined over two categorical axes (time × category). Color encodes the quantity.
Pair plot: Exploratory analysis of up to ~10 variables. Shows all pairwise scatter plots simultaneously for rapid pattern detection.
- Scatter plots reveal direction, form, strength, and unusual features of the relationship between two quantitative variables.
- Pearson’s r quantifies the strength of a linear relationship; Spearman’s ρ is more robust for nonlinear or rank-ordered data.
- Correlation does not imply causation — always consider whether a third variable could explain the observed association.
- Line plots are the natural choice when one axis is time and temporal order matters.
- Bubble charts add a third dimension via point size; heatmaps handle many variables via color; pair plots give a simultaneous overview of all pairwise relationships.