STAT 101
M7 · Lab
Hands-on lab
Fit a line, then find its influential point

Fit a least-squares line in the regression sandbox, and predict — before you read it off the screen — the slope, the intercept, the R², and then the leverage and Cook's distance of a far-out point that quietly swings the whole line.

Cook's distance and leverage — P counts PARAMETERS, not predictors
D_i = \frac{e_i^{2}}{P\,s^{2}}\cdot\frac{h_{ii}}{(1-h_{ii})^{2}}, \quad h_{ii} = \frac{1}{n} + \frac{(x_i-\bar{x})^{2}}{S_{xx}}

A high R² tells you the line fits the points you have — it does not tell you a single point is not dragging the whole fit behind it. Leverage is how far a point's x sits from the others; Cook's distance is how much the fitted line would move if you dropped the point. A point can have high leverage and almost no influence, or a modest leverage and enormous influence — and reading the two together is how you catch the one observation that is writing your conclusion for you.

1 / 9
STAT 101
M7 · Lab
Set it up
One fit

Open the sandbox. It opens on exactly this state — the Clean linear preset, noise seed 3 — so a Reset gets you here too. Each step names the one knob to change; leave everything else alone.

Set these values
Dataset -> Clean linear Noise seed -> 3 Point -> 11 (the last) Bands -> none

Watch the slope, intercept and R² readouts, and — in panel 4 — the per-point leverage and Cook's distance with their thresholds. Every number you predict is printed there.

2 / 9
STAT 101
M7 · Lab
Step 1 of 5
The least-squares slope, b₁

Keep the defaults — the Clean linear preset, noise seed 3. The least-squares line is the one that minimises the sum of squared residuals. The slope b₁ is the change in ŷ per unit x. Predict it to two decimals, then read the slope readout.

Expected

The slope readout is b₁ = 1.2508 — each unit increase in x raises the fitted ŷ by about 1.25. This fit has n = 12 points and k = 1 predictor.

3 / 9
STAT 101
M7 · Lab
Step 2 of 5
The intercept, b₀

Same fit. The intercept b₀ is the fitted ŷ where x = 0. The least-squares line always passes through the point of means (x̄, ȳ), so b₀ = ȳ − b₁·x̄. Predict b₀, then read the intercept readout.

Expected

The intercept readout is b₀ = 3.0791 — the line crosses the y-axis near 3.08. It is an extrapolation here: the data start at x = 1, so x = 0 lies just outside the observed range.

4 / 9
STAT 101
M7 · Lab
Step 3 of 5
How much variance does R² explain?

Still the clean fit. R² is the fraction of the variance in y the line explains: R² = 1 − SSE/SST. Predict it, then read the R² readout.

Expected

The readout shows R² = 0.9429 — the line explains about 94% of the variance in y. But a high R² alone never proves the model is right: the residual plot still has to be flat and patternless, which is what the next steps stress-test.

5 / 9
STAT 101
M7 · Lab
Step 4 of 5
The leverage of a far point, hᵢᵢ

Switch Dataset to High leverage, OFF the trend, keeping noise seed 3 and point 11 selected — it sits far out at x = 22. Leverage hᵢᵢ = 1/n + (xᵢ − x̄)²/Sₓₓ depends on x alone. The rule of thumb flags hᵢᵢ > 2P/n. Predict whether point 11 is flagged, then read its leverage in panel 4.

Expected

The leverage is h₁₁ = 0.8075 — far above the cutoff. With n = 12 and P = 2 PARAMETERS (k = 1 predictor plus the intercept), that cutoff is 2P/n = 0.333, so point 11 is flagged. The leverages sum to P = 2 — which is exactly why the cutoff carries P, not k.

6 / 9
STAT 101
M7 · Lab
Step 5 of 5
Cook's distance — the influence

Same point 11 on the same preset. Cook's distance combines the residual and the leverage, dividing by P·s² — P = 2 PARAMETERS, not k = 1 predictor (dividing by k instead would double it). Cook flags D > 1. Predict whether point 11 is influential, then read its Cook's D.

Expected

Cook's distance is D₁₁ = 19.6268 — far past the D > 1 rule (Cook 1977). This far point is not merely high-leverage, it is genuinely influential: drop it and the whole line swings. High leverage with a small residual is harmless; high leverage plus a big residual is what moves the fit.

7 / 9
STAT 101
M7 · Lab
Your turn
Open the sandbox

Everything above is waiting in the sandbox. Drag any point and watch the fitted line chase it, step through the presets to see a curved fit, a fan of growing variance and a mid-range outlier, and select any observation to read its leverage, studentised residual and Cook's distance side by side.

8 / 9
STAT 101
M7 · Lab
Wrap-up
What you did
  • Read the least-squares slope b₁ = 1.2508 and intercept b₀ = 3.0791
  • Confirmed the line explains most of the variance, R² = 0.9429
  • Flagged a far point for leverage, h₁₁ = 0.8075 > 2P/n = 0.333 (n = 12, P = 2)
  • Measured its influence, Cook's D₁₁ = 19.6268 > 1 — high leverage AND a big residual
  • Every value you predicted is the demo's own regression arithmetic, not a picture — and Cook's D divides by P, not k
9 / 9