Why Diagnostics Matter
Fitting a regression model is easy. Trusting its output is a different matter. Ordinary Least Squares (OLS) guarantees the Best Linear Unbiased Estimator (BLUE) only when a set of assumptions hold: linearity in the parameters, zero-mean errors, constant variance (homoscedasticity), independence of errors, and, for inference, approximate normality of residuals. When these assumptions fail, coefficient estimates can be biased, standard errors can be wrong, p-values misleading, and prediction intervals too narrow or too wide.
Regression diagnostics are the toolkit for checking whether these assumptions are met. Rather than guessing, we examine the residuals — the differences between observed and fitted values — and look for patterns that signal violations. Diagnostics also reveal influential observations that disproportionately shape the estimated model. This lesson covers the most important diagnostic plots and statistics, how to interpret them, and what to do when they flag a problem.
Residuals: The Diagnostic Signal
The raw residual for observation i is simply:
It is often more useful to work with standardized residuals, which divide the raw residual by an estimate of its standard deviation:
Studentized (externally studentized) residuals go one step further: they refit the model with observation i deleted and compute the residual relative to that leave-one-out estimate. Under the null model, studentized residuals follow a t-distribution with n−p−2 degrees of freedom, enabling formal outlier tests (the Bonferroni-corrected version is called the Bonferroni outlier test).
Residual Plots: Reading the Patterns
The most informative diagnostic is the residuals vs. fitted values plot. Under a correctly specified model the plot should look like a horizontal band of random scatter around zero. Common problematic patterns are:
Curved band (U-shape or arch): non-linearity — the model is missing a quadratic or other nonlinear term. Add polynomial features or transform the predictor.
Fan shape (spread increases with fitted values): heteroscedasticity — the error variance is not constant. Use weighted least squares, a variance-stabilizing transformation (log Y), or robust standard errors.
Obvious outliers (points far from the band): potential influential observations — investigate whether they are data errors, exceptional cases, or genuinely unusual but valid data points.
A secondary useful plot is residuals vs. each predictor. If a predictor’s residual plot shows a curve, that predictor might need a transformation or a squared term. The scale-location plot (square root of |standardized residuals| vs. fitted values) makes heteroscedasticity easier to see as an upward slope in a smooth line fit through the points.
Testing Normality: The QQ Plot
OLS estimates are unbiased and consistent regardless of whether errors are normal. But the t-tests and F-tests for significance, and the prediction intervals, rely on normality. For small to moderate samples the Normal QQ plot (quantile-quantile plot) is the standard check.
A QQ plot sorts the standardized residuals from smallest to largest and plots them against the corresponding quantiles of the standard normal distribution. If residuals are truly normal, the points fall on a straight 45° line. Common deviations and their meanings:
S-curve (tails above then below line): heavy tails (leptokurtosis) — the residuals have more extreme values than a normal distribution predicts. Common in financial data. Bootstrapped inference or robust methods help.
Inverted S-curve: light tails (platykurtosis) — relatively rare in practice, usually benign.
Points curving up at both ends: right-skewed distribution — consider a log transformation of the response.
One or two isolated points far from the line: potential outliers — investigate individually.
For formal testing, the Shapiro-Wilk test is powerful for small samples (n ≤ 50). For larger samples the Anderson-Darling or Kolmogorov-Smirnov tests are used, but any departure from normality becomes statistically significant with large enough n — the QQ plot remains the most informative tool even then, because it shows the nature and magnitude of departures.
Homoscedasticity: Constant Error Variance
OLS standard errors are derived under the assumption that Var(εᵢ) = σ² for all i. When error variance depends on the predictors or the fitted values — a violation called heteroscedasticity — OLS coefficient estimates remain unbiased but their standard errors are wrong, making hypothesis tests unreliable.
The most common formal test is the Breusch-Pagan test: regress the squared residuals on the original predictors and check whether the slope coefficients are jointly significant (using an F-test or chi-squared test). The null hypothesis is homoscedasticity; a significant result means variance is not constant. The White test is a more general version that also includes cross-product terms.
Variance-stabilizing transformation: if Y is right-skewed (e.g., income, house prices), try log(Y) as the response. Log transformation often equates variance across the range of fitted values because multiplicative percentage errors become additive on the log scale.
Weighted Least Squares (WLS): if the variance structure is known (e.g., Var(εᵢ) ∝ xᵢ), assign observation weights wᵢ = 1/Var(εᵢ). WLS recovers efficiency, but the variance model must be correctly specified.
Heteroscedasticity-Consistent (HC) standard errors: also called “robust” or “sandwich” standard errors, these require no assumption about the form of heteroscedasticity. They are the default in econometrics and are discussed in the last section of this lesson.
Influential Points: Leverage
Not every outlier harms the model, and not every unusual predictor value (high-leverage point) is harmful either. The danger arises when a point is both unusual in its X values and has a large residual — such points can pull the regression line strongly toward themselves.
The leverage hᵢᵢ of observation i is the i-th diagonal element of the hat matrix H = X(X′X)⁻¹X′. It measures how far observation i’s predictor values are from the centroid of all predictor values:
High leverage alone is not necessarily harmful. A point at an extreme X value with a residual consistent with the model fits perfectly — it is even beneficial in precisely estimating the slope in that region. Problems arise only when high leverage combines with a large residual.
Cook’s Distance
Cook’s distance D₁ combines leverage and residual size into a single measure of how much the coefficient vector changes when observation i is removed:
Common thresholds: Cook’s D > 1 is a traditional “highly influential” threshold. A modern approach plots Dᵢ against observation index and looks for points that stand out from the rest — even D < 1 can be concerning if one observation is far above all others. DFFITS (difference in fits) and DFBETAS (difference in each β) are related measures that can pinpoint which coefficient is most affected.
Never delete an influential observation automatically. First verify it is not a data entry or recording error. If it is valid, investigate why it deviates — it may reveal genuine heterogeneity in the population (e.g., a subgroup that follows a different relationship). Options include: fit the model with and without the observation and report both; use robust regression methods (M-estimators, median regression) that downweight extreme observations; or model the subgroup separately.
Transformations
Many real-world relationships are nonlinear, and many response variables have non-normal, heteroscedastic errors in their raw scale. A well-chosen transformation can simultaneously linearize the relationship and stabilize variance, allowing OLS to be applied validly.
The log transformation is the single most useful transformation in regression. When Y is positive and right-skewed (incomes, prices, concentrations), taking log(Y) as the response typically produces residuals that are more nearly normal and homoscedastic. The model log(Y) = β₀ + β₁X implies that a one-unit increase in X multiplies Y by e^(β₁) — a multiplicative rather than additive effect.
The Box-Cox family of power transformations generalizes log and square root into a single parameterized family:
Transforming the predictors is also useful. If a scatterplot of Y vs. X shows a curved relationship, adding X² (polynomial regression) or taking log(X) can linearize it. For count predictors bounded at zero, log(1 + X) avoids the log(0) problem. Be mindful that transforming the response changes the interpretation of all coefficients and makes back-transforming predictions slightly complex (the geometric mean of Y is predicted, not the arithmetic mean).
Heteroscedasticity-Robust Standard Errors
When heteroscedasticity is present but you do not want to transform the response or impose a variance model, heteroscedasticity-consistent (HC) standard errors — also called “sandwich” or “robust” standard errors — provide valid inference without requiring that Var(εᵢ) be constant.
The idea: the standard OLS variance-covariance matrix (X′X)⁻¹σ² assumes each squared error εᵢ² equals the same σ². The HC estimator replaces σ² with the actual squared residual eᵢ² for each observation, then averages appropriately:
Robust standard errors do not change the point estimates β̂ — only the standard errors (and thus t-statistics and confidence intervals). They are widely used in econometrics, political science, and public health. In R, use coeftest(..., vcov = vcovHC(...)) from the sandwich package; in Python, use fit().get_robustcov_results(cov_type='HC3') from statsmodels.
- Residuals vs. fitted is the primary diagnostic plot: random scatter = good; curves = non-linearity; fan shape = heteroscedasticity; isolated outliers = investigate.
- QQ plot checks residual normality; an S-curve suggests heavy tails, a one-sided curve suggests skew — consider transforming the response.
- Heteroscedasticity biases standard errors but not estimates; remedies include log transformation, weighted least squares, or robust (HC) standard errors.
- Leverage hᵢᵢ measures how extreme an observation is in predictor space; high-leverage observations exert disproportionate influence on the fitted line.
- Cook’s distance combines leverage and residual to quantify total influence; investigate outliers with D > 1 or those that stand far above the rest.
- Box-Cox transformations systematically find the power λ that best normalizes residuals and stabilizes variance; log (λ = 0) is the most common choice.
- HC sandwich standard errors give valid t-tests under heteroscedasticity without changing coefficient estimates; they are the default robust option in econometrics.