From One Predictor to Many
Real phenomena are rarely driven by a single variable. House prices depend on floor area, location, number of bedrooms, age, and dozens of other factors simultaneously. Multiple regression extends the simple linear model to accommodate multiple predictors, each contributing independently to the predicted outcome.
The jump from one predictor to many is not merely a notational change — it introduces new concepts (partial effects, multicollinearity, adjusted R²) and new challenges (overfitting, variable selection, categorical encoding). Mastering these ideas is essential before applying any regression model to real data.
The Multiple Regression Model
With p predictors X₁, X₂, …, Xᴝ the model is:
In matrix notation this becomes Y = Xβ + ε, where Y is an n×1 vector of responses, X is an n×(p+1) design matrix (first column all ones for the intercept), β is a (p+1)×1 coefficient vector, and ε is an n×1 error vector. The OLS estimator is:
Partial Regression Coefficients
The coefficient β̂₁ in simple regression estimates how Y changes per unit of X₁, ignoring all other variables. In multiple regression, β̂₁ is a partial regression coefficient — it estimates how Y changes per unit of X₁ while holding all other predictors constant.
This “holding constant” interpretation is crucial. Suppose we regress salary on years of experience and education level. The coefficient on experience now tells us the salary premium per year of experience for people at the same education level — not the raw correlation between experience and salary, which is inflated by the tendency of more-educated people to also have different experience profiles.
We model house price (in $k) as: Price = −18.4 + 0.09 × Area + 32.1 × Bedrooms + −0.7 × Age. The coefficient 0.09 means each additional square foot adds $90 to the predicted price for houses with the same number of bedrooms and the same age. The raw correlation between area and price is higher because larger houses also tend to have more bedrooms.
Adjusted R²
Adding any predictor to a regression model — even a random noise variable — cannot decrease R². This is a mathematical consequence of OLS: more parameters can always fit the existing data at least as well. A naive researcher could inflate R² simply by adding irrelevant variables.
Adjusted R² corrects for this by penalizing the number of predictors:
Adding a predictor increases adjusted R² only if its t-statistic exceeds 1 in absolute value — a mild hurdle, but it filters out pure noise. In practice, use adjusted R² when comparing models with different numbers of predictors, never plain R².
Multicollinearity
Multicollinearity occurs when two or more predictors are highly correlated with each other. It does not bias the coefficient estimates, but it inflates their standard errors, making individual coefficients imprecise and statistically insignificant even when the overall model fits well.
Consider regressing blood pressure on age, weight, and body-mass-index (BMI). Since BMI = weight / height², BMI is nearly a function of weight. Including both means the coefficient on weight is trying to capture the unique effect of weight after controlling for something that is almost entirely determined by weight — an ill-posed problem.
The standard diagnostic is the Variance Inflation Factor (VIF):
Remedies include: removing one of the correlated predictors (domain knowledge helps choose which), combining correlated predictors into an index or composite score, using Ridge regression (which shrinks coefficients and tolerates collinearity), or collecting more data (which reduces standard errors but does not eliminate collinearity).
Feature Selection
With p candidate predictors, 2ᴝ possible subsets exist. Searching all of them (“best subset selection”) becomes computationally infeasible once p exceeds roughly 30–40. In practice, greedy sequential algorithms are used:
Forward selection starts with no predictors and adds one at a time, at each step choosing the predictor that most improves a criterion (AIC, BIC, adjusted R², or the smallest p-value). It stops when no remaining variable passes the threshold.
Backward elimination starts with all p predictors and removes one at a time, discarding whichever contributes least (largest p-value, smallest F-to-remove). It stops when all remaining variables pass the threshold.
Stepwise (bidirectional) alternates forward and backward steps, allowing variables entered in earlier steps to be removed if they become insignificant after new variables are added, and vice versa. It is more thorough than either unidirectional method but still not guaranteed to find the global optimum.
Stepwise methods select variables based on the same data used to estimate them, inflating Type I error rates and producing overly optimistic model statistics. They should be treated as exploratory tools, not confirmatory ones. Whenever possible, use theory or prior knowledge to guide variable inclusion, or use penalized regression methods (Lasso, Ridge) which simultaneously estimate and select variables in a statistically principled way.
Categorical Predictors: Dummy Variables
Regression requires numeric inputs, but many important predictors are categorical: region (North, South, East, West), treatment group (A, B, Control), or education level (High School, Bachelor’s, Graduate). These are incorporated via dummy variables (also called indicator variables).
For a categorical variable with k levels, we create k−1 binary dummy variables. One level is designated the reference category (omitted to avoid perfect multicollinearity with the intercept — the “dummy variable trap”). Each dummy coefficient is interpreted as the effect of that category relative to the reference category.
Regressing sales on advertising spend and region (North = reference, South, East, West), we include three dummies: D₁ = 1 if South, D₂ = 1 if East, D₃ = 1 if West. The coefficient on D₁ is the average sales difference between South and North after controlling for advertising spend. The choice of reference category affects the coefficient values but not the fitted values, R², or any overall model statistics.
Interaction terms between a dummy and a continuous predictor allow the slope to differ across categories. For example, D₁ × Advertising spend would capture whether the return on advertising differs between South and North. Without the interaction, the model assumes the same advertising slope in all regions; with it, each region has its own slope.
Model Comparison: AIC and BIC
When comparing models with different numbers of predictors, adjusted R² is one option, but information criteria — AIC (Akaike) and BIC (Bayesian/Schwarz) — are widely preferred because they have stronger theoretical grounding and naturally balance fit against complexity:
AIC is preferred when the goal is prediction; BIC is preferred when the goal is identifying the true model structure. For typical regression models, k = p + 2 (p slopes plus the intercept and σ²). When n is large, BIC’s heavier penalty means it will select sparser models than AIC.
- Multiple regression Y = β₀ + β₁X₁ + … + βᴝXᴝ + ε generalizes simple regression to p predictors with the OLS solution β̂ = (X′X)¹X′Y.
- Partial coefficients measure the effect of each predictor while holding all others fixed — they differ from marginal correlations and require careful interpretation.
- Adjusted R² penalizes model complexity so it increases only when a new predictor genuinely improves fit beyond chance.
- Multicollinearity inflates standard errors and destabilizes individual coefficients; diagnose with VIF (flag if > 5–10) and remediate by removing correlated variables or using Ridge regression.
- Feature selection via forward, backward, or stepwise procedures is exploratory — it overstates significance when used confirmatorily. Prefer theory-driven or penalized approaches.
- Categorical predictors enter via k−1 dummy variables with one reference level; interaction dummies allow slopes to vary across categories.
- AIC and BIC compare models with different numbers of parameters; lower is better; BIC favors parsimony more strongly for large n.