Reading
Stories Mode

Multiple Regression

~30 min read Lesson 2 of 4 in Module 7

From One Predictor to Many

Real phenomena are rarely driven by a single variable. House prices depend on floor area, location, number of bedrooms, age, and dozens of other factors simultaneously. Multiple regression extends the simple linear model to accommodate multiple predictors, each contributing independently to the predicted outcome.

The jump from one predictor to many is not merely a notational change — it introduces new concepts (partial effects, multicollinearity, adjusted R²) and new challenges (overfitting, variable selection, categorical encoding). Mastering these ideas is essential before applying any regression model to real data.

The Multiple Regression Model

With p predictors X₁, X₂, …, Xᴝ the model is:

Multiple Regression Model
Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_p X_p + \varepsilon
Y is the response; β₀ the intercept; β₁,…,βₚ the partial regression coefficients; X₁,…,Xₚ the predictors; ε the error term with mean zero and constant variance σ².

In matrix notation this becomes Y = Xβ + ε, where Y is an n×1 vector of responses, X is an n×(p+1) design matrix (first column all ones for the intercept), β is a (p+1)×1 coefficient vector, and ε is an n×1 error vector. The OLS estimator is:

OLS Estimator (Matrix Form)
\hat{\boldsymbol{\beta}} = (\mathbf{X}^\prime \mathbf{X})^{-1} \mathbf{X}^\prime \mathbf{Y}
This closed-form solution requires X'X to be invertible — guaranteed when no predictor is a perfect linear combination of others (no perfect multicollinearity).

Partial Regression Coefficients

The coefficient β̂₁ in simple regression estimates how Y changes per unit of X₁, ignoring all other variables. In multiple regression, β̂₁ is a partial regression coefficient — it estimates how Y changes per unit of X₁ while holding all other predictors constant.

This “holding constant” interpretation is crucial. Suppose we regress salary on years of experience and education level. The coefficient on experience now tells us the salary premium per year of experience for people at the same education level — not the raw correlation between experience and salary, which is inflated by the tendency of more-educated people to also have different experience profiles.

Worked Example

We model house price (in $k) as: Price = −18.4 + 0.09 × Area + 32.1 × Bedrooms + −0.7 × Age. The coefficient 0.09 means each additional square foot adds $90 to the predicted price for houses with the same number of bedrooms and the same age. The raw correlation between area and price is higher because larger houses also tend to have more bedrooms.

Adjusted R²

Adding any predictor to a regression model — even a random noise variable — cannot decrease R². This is a mathematical consequence of OLS: more parameters can always fit the existing data at least as well. A naive researcher could inflate R² simply by adding irrelevant variables.

Adjusted R² corrects for this by penalizing the number of predictors:

Adjusted R²
\bar{R}^2 = 1 - \dfrac{\text{SSR}/(n-p-1)}{\text{SST}/(n-1)} = 1 - (1 - R^2)\dfrac{n-1}{n-p-1}
n = sample size, p = number of predictors (not counting the intercept). Adjusted R² increases only if the new predictor improves fit more than expected by chance. It can be negative for very poor models.

Adding a predictor increases adjusted R² only if its t-statistic exceeds 1 in absolute value — a mild hurdle, but it filters out pure noise. In practice, use adjusted R² when comparing models with different numbers of predictors, never plain R².

Multicollinearity

Multicollinearity occurs when two or more predictors are highly correlated with each other. It does not bias the coefficient estimates, but it inflates their standard errors, making individual coefficients imprecise and statistically insignificant even when the overall model fits well.

Consider regressing blood pressure on age, weight, and body-mass-index (BMI). Since BMI = weight / height², BMI is nearly a function of weight. Including both means the coefficient on weight is trying to capture the unique effect of weight after controlling for something that is almost entirely determined by weight — an ill-posed problem.

The standard diagnostic is the Variance Inflation Factor (VIF):

Variance Inflation Factor
\text{VIF}_j = \dfrac{1}{1 - R_j^2}
Rⱼ² is the R² from regressing Xⱼ on all other predictors. VIF = 1 means no multicollinearity; VIF > 5 is a common warning threshold; VIF > 10 indicates severe multicollinearity requiring action.

Remedies include: removing one of the correlated predictors (domain knowledge helps choose which), combining correlated predictors into an index or composite score, using Ridge regression (which shrinks coefficients and tolerates collinearity), or collecting more data (which reduces standard errors but does not eliminate collinearity).

Feature Selection

With p candidate predictors, 2ᴝ possible subsets exist. Searching all of them (“best subset selection”) becomes computationally infeasible once p exceeds roughly 30–40. In practice, greedy sequential algorithms are used:

Forward selection starts with no predictors and adds one at a time, at each step choosing the predictor that most improves a criterion (AIC, BIC, adjusted R², or the smallest p-value). It stops when no remaining variable passes the threshold.

Backward elimination starts with all p predictors and removes one at a time, discarding whichever contributes least (largest p-value, smallest F-to-remove). It stops when all remaining variables pass the threshold.

Stepwise (bidirectional) alternates forward and backward steps, allowing variables entered in earlier steps to be removed if they become insignificant after new variables are added, and vice versa. It is more thorough than either unidirectional method but still not guaranteed to find the global optimum.

Limitations of Stepwise Procedures

Stepwise methods select variables based on the same data used to estimate them, inflating Type I error rates and producing overly optimistic model statistics. They should be treated as exploratory tools, not confirmatory ones. Whenever possible, use theory or prior knowledge to guide variable inclusion, or use penalized regression methods (Lasso, Ridge) which simultaneously estimate and select variables in a statistically principled way.

Categorical Predictors: Dummy Variables

Regression requires numeric inputs, but many important predictors are categorical: region (North, South, East, West), treatment group (A, B, Control), or education level (High School, Bachelor’s, Graduate). These are incorporated via dummy variables (also called indicator variables).

For a categorical variable with k levels, we create k−1 binary dummy variables. One level is designated the reference category (omitted to avoid perfect multicollinearity with the intercept — the “dummy variable trap”). Each dummy coefficient is interpreted as the effect of that category relative to the reference category.

Example: Regional Effects

Regressing sales on advertising spend and region (North = reference, South, East, West), we include three dummies: D₁ = 1 if South, D₂ = 1 if East, D₃ = 1 if West. The coefficient on D₁ is the average sales difference between South and North after controlling for advertising spend. The choice of reference category affects the coefficient values but not the fitted values, R², or any overall model statistics.

Interaction terms between a dummy and a continuous predictor allow the slope to differ across categories. For example, D₁ × Advertising spend would capture whether the return on advertising differs between South and North. Without the interaction, the model assumes the same advertising slope in all regions; with it, each region has its own slope.

Model Comparison: AIC and BIC

When comparing models with different numbers of predictors, adjusted R² is one option, but information criteria — AIC (Akaike) and BIC (Bayesian/Schwarz) — are widely preferred because they have stronger theoretical grounding and naturally balance fit against complexity:

AIC and BIC
\text{AIC} = -2\ell + 2k \qquad \text{BIC} = -2\ell + k\ln n
ℓ is the maximized log-likelihood, k is the number of estimated parameters (including σ²), and n is the sample size. Lower AIC or BIC = better model. BIC penalizes complexity more heavily than AIC when n > 7, making it prefer simpler models.

AIC is preferred when the goal is prediction; BIC is preferred when the goal is identifying the true model structure. For typical regression models, k = p + 2 (p slopes plus the intercept and σ²). When n is large, BIC’s heavier penalty means it will select sparser models than AIC.

Key Takeaways
  • Multiple regression Y = β₀ + β₁X₁ + … + βᴝXᴝ + ε generalizes simple regression to p predictors with the OLS solution β̂ = (X′X)¹X′Y.
  • Partial coefficients measure the effect of each predictor while holding all others fixed — they differ from marginal correlations and require careful interpretation.
  • Adjusted R² penalizes model complexity so it increases only when a new predictor genuinely improves fit beyond chance.
  • Multicollinearity inflates standard errors and destabilizes individual coefficients; diagnose with VIF (flag if > 5–10) and remediate by removing correlated variables or using Ridge regression.
  • Feature selection via forward, backward, or stepwise procedures is exploratory — it overstates significance when used confirmatorily. Prefer theory-driven or penalized approaches.
  • Categorical predictors enter via k−1 dummy variables with one reference level; interaction dummies allow slopes to vary across categories.
  • AIC and BIC compare models with different numbers of parameters; lower is better; BIC favors parsimony more strongly for large n.
Previous Simple Linear Regression Module Overview Next Regression Diagnostics