STAT 101
M07 · L03
Module 7 · Lesson 3
Regression Diagnostics
Fitting a model is just the beginning. Diagnostics tell you whether the model’s assumptions hold — and whether any observations are quietly distorting your results.
01 / 11
STAT 101
M07 · L03
The diagnostic signal
The Residual
Raw Residual
e_i = y_i - \hat{y}_i
Residuals should scatter randomly around zero. Any pattern — curve, fan, cluster — signals a model problem.
02 / 11
STAT 101
M07 · L03
Residuals vs. fitted values
Reading Patterns
- Random band — assumptions met ✓
- Curved (U-shape) — non-linearity; add polynomial term
- Fan shape — heteroscedasticity; variance increases
- Isolated outlier — investigate for data error or influence
03 / 11
STAT 101
M07 · L03
Checking normality
The QQ Plot
How to read it
Straight line = normal residuals ✓
S-curve = heavy tails (kurtosis)
One-sided curve = skew → try log(Y)
Single stray point = outlier
S-curve = heavy tails (kurtosis)
One-sided curve = skew → try log(Y)
Single stray point = outlier
OLS inference relies on approximate normality. The QQ plot is more informative than any formal test at revealing the nature of departures.
04 / 11
STAT 101
M07 · L03
Non-constant variance
Heteroscedasticity
- Effect: standard errors are wrong — t-tests invalid
- Test: Breusch-Pagan: regress eᵢ² on predictors
- Fix 1: transform Y (e.g., log)
- Fix 2: weighted least squares
- Fix 3: robust (HC) standard errors
05 / 11
STAT 101
M07 · L03
Extreme predictor values
Leverage hii
Hat value
h_{ii} = \mathbf{x}_i^\prime (\mathbf{X}^\prime\mathbf{X})^{-1} \mathbf{x}_i
Range
1/n to 1
Flag if
> 2(p+1)/n
06 / 11
STAT 101
M07 · L03
Combined influence measure
Cook’s Distance
Cook's D
D_i = \frac{e_i^2}{(p+1)\,s^2} \cdot \frac{h_{ii}}{(1-h_{ii})^2}
Dᵢ > 1: traditionally “highly influential.” More practically, look for points that stand far above the others in a Cook’s D vs. index plot.
07 / 11
STAT 101
M07 · L03
Linearizing & stabilizing variance
Box-Cox Family
Power transformation
y^{(\lambda)} = \begin{cases} (y^\lambda - 1)/\lambda & \lambda \neq 0 \\ \ln y & \lambda = 0 \end{cases}
λ estimated by MLE. λ = 0 gives log(Y); λ = 0.5 gives √Y. A 95% CI for λ guides practical choice.
08 / 11
STAT 101
M07 · L03
Valid inference under heteroscedasticity
Sandwich Errors
HC3 robust SE
Uses each observation’s own residual eᵢ
No assumption on error variance form
Slightly larger than OLS SE when homoscedastic
Correct when heteroscedastic
No assumption on error variance form
Slightly larger than OLS SE when homoscedastic
Correct when heteroscedastic
09 / 11
STAT 101
Summary
Recap
What you learned
- Residuals vs. fitted: random band = OK; curves or fans = problems
- QQ plot checks normality; S-curve = heavy tails, slope = skew
- Heteroscedasticity invalidates SE; fix with log, WLS, or HC errors
- Leverage hᵢᵢ: how far X values are from the centroid
- Cook’s D: combined leverage + residual — total influence
- Box-Cox (λ) finds the best power transformation for Y
11 / 11