Fit a least-squares line in the regression sandbox, and predict — before you read it off the screen — the slope, the intercept, the R², and then the leverage and Cook's distance of a far-out point that quietly swings the whole line.
A high R² tells you the line fits the points you have — it does not tell you a single point is not dragging the whole fit behind it. Leverage is how far a point's x sits from the others; Cook's distance is how much the fitted line would move if you dropped the point. A point can have high leverage and almost no influence, or a modest leverage and enormous influence — and reading the two together is how you catch the one observation that is writing your conclusion for you.
Open the sandbox. It opens on exactly this state — the Clean linear preset, noise seed 3 — so a Reset gets you here too. Each step names the one knob to change; leave everything else alone.
Dataset -> Clean linear Noise seed -> 3
Point -> 11 (the last) Bands -> none
Watch the slope, intercept and R² readouts, and — in panel 4 — the per-point leverage and Cook's distance with their thresholds. Every number you predict is printed there.
Keep the defaults — the Clean linear preset, noise seed 3. The least-squares line is the one that minimises the sum of squared residuals. The slope b₁ is the change in ŷ per unit x. Predict it to two decimals, then read the slope readout.
The slope readout is b₁ = 1.2508 — each unit increase in x raises the fitted ŷ by about 1.25. This fit has n = 12 points and k = 1 predictor.
Same fit. The intercept b₀ is the fitted ŷ where x = 0. The least-squares line always passes through the point of means (x̄, ȳ), so b₀ = ȳ − b₁·x̄. Predict b₀, then read the intercept readout.
The intercept readout is b₀ = 3.0791 — the line crosses the y-axis near 3.08. It is an extrapolation here: the data start at x = 1, so x = 0 lies just outside the observed range.
Still the clean fit. R² is the fraction of the variance in y the line explains: R² = 1 − SSE/SST. Predict it, then read the R² readout.
The readout shows R² = 0.9429 — the line explains about 94% of the variance in y. But a high R² alone never proves the model is right: the residual plot still has to be flat and patternless, which is what the next steps stress-test.
Switch Dataset to High leverage, OFF the trend, keeping noise seed 3 and point 11 selected — it sits far out at x = 22. Leverage hᵢᵢ = 1/n + (xᵢ − x̄)²/Sₓₓ depends on x alone. The rule of thumb flags hᵢᵢ > 2P/n. Predict whether point 11 is flagged, then read its leverage in panel 4.
The leverage is h₁₁ = 0.8075 — far above the cutoff. With n = 12 and P = 2 PARAMETERS (k = 1 predictor plus the intercept), that cutoff is 2P/n = 0.333, so point 11 is flagged. The leverages sum to P = 2 — which is exactly why the cutoff carries P, not k.
Same point 11 on the same preset. Cook's distance combines the residual and the leverage, dividing by P·s² — P = 2 PARAMETERS, not k = 1 predictor (dividing by k instead would double it). Cook flags D > 1. Predict whether point 11 is influential, then read its Cook's D.
Cook's distance is D₁₁ = 19.6268 — far past the D > 1 rule (Cook 1977). This far point is not merely high-leverage, it is genuinely influential: drop it and the whole line swings. High leverage with a small residual is harmless; high leverage plus a big residual is what moves the fit.
Everything above is waiting in the sandbox. Drag any point and watch the fitted line chase it, step through the presets to see a curved fit, a fan of growing variance and a mid-range outlier, and select any observation to read its leverage, studentised residual and Cook's distance side by side.
- Read the least-squares slope b₁ = 1.2508 and intercept b₀ = 3.0791
- Confirmed the line explains most of the variance, R² = 0.9429
- Flagged a far point for leverage, h₁₁ = 0.8075 > 2P/n = 0.333 (n = 12, P = 2)
- Measured its influence, Cook's D₁₁ = 19.6268 > 1 — high leverage AND a big residual
- Every value you predicted is the demo's own regression arithmetic, not a picture — and Cook's D divides by P, not k