STAT 101
M07 · L02
Module 7 · Lesson 2
Multiple Regression
Real outcomes depend on many factors at once. Multiple regression extends the linear model to p predictors — estimating each variable’s unique contribution while controlling for all the others.
01 / 12
STAT 101
M07 · L02
The equation
The Model
Multiple Linear Regression
Y = \beta_0 + \beta_1 X_1 + \cdots + \beta_p X_p + \varepsilon
In matrix form: Y = Xβ + ε. The OLS solution is β̂ = (X′X)⁻¹X′Y — a single formula for any number of predictors.
02 / 12
STAT 101
M07 · L02
Holding others fixed
Partial Effects
Key distinction
β̂₁ in simple regression: raw change in Y per unit X₁
β̂₁ in multiple regression: change in Y per unit X₁ with X₂,…,Xᴝ held constant
β̂₁ in multiple regression: change in Y per unit X₁ with X₂,…,Xᴝ held constant
Partial coefficients control for confounders — they measure the unique linear contribution of each predictor.
03 / 12
STAT 101
M07 · L02
Penalized goodness of fit
Adjusted R²
Adjusted R-squared
\bar{R}^2 = 1 - (1-R^2)\dfrac{n-1}{n-p-1}
Plain R² never decreases when predictors are added. Adjusted R² rises only if the new variable genuinely helps — it penalizes model complexity.
04 / 12
STAT 101
M07 · L02
Correlated predictors
Multicollinearity
- What it is: two or more predictors highly correlated with each other
- Effect: inflated standard errors — individual coefficients become unstable
- Sign: overall model fits well (high R²) but no predictor is significant
- Fix: remove one variable, combine, or use Ridge regression
05 / 12
STAT 101
M07 · L02
Diagnosing collinearity
Variance Inflation Factor
VIF
\text{VIF}_j = \dfrac{1}{1 - R_j^2}
VIF = 1
No collinearity
VIF > 10
Severe — act
06 / 12
STAT 101
M07 · L02
Which predictors to keep?
Feature Selection
- Forward: add the best predictor one at a time
- Backward: start full, remove the worst one at a time
- Stepwise: forward + backward interleaved
- Limitation: all inflate Type I error — treat as exploratory
07 / 12
STAT 101
M07 · L02
Categorical predictors
Dummy Variables
k levels → k−1 dummies
Region: North (reference), South, East, West
→ 3 dummies: D₁, D₂, D₃
Each coefficient = effect vs. reference level, holding other predictors constant
→ 3 dummies: D₁, D₂, D₃
Each coefficient = effect vs. reference level, holding other predictors constant
Adding a dummy × continuous interaction lets the slope differ across categories.
08 / 12
STAT 101
M07 · L02
Model comparison
AIC & BIC
Information Criteria
\text{AIC} = {-2\ell + 2k} \quad \text{BIC} = {-2\ell + k\ln n}
AIC
Favors prediction
BIC
Favors parsimony
09 / 12
STAT 101
M07 · L02
Same assumptions as simple regression
Checking Assumptions
- Linearity — residuals vs. fitted: no curve
- Homoscedasticity — constant spread in residuals
- Normality — QQ plot of residuals
- Independence — no autocorrelation
- No multicollinearity — check VIF for each predictor
10 / 12
STAT 101
Summary
Recap
What you learned
- Model: Y = β₀ + β₁X₁ + … + βᴝXᴝ + ε; solved by β̂ = (X′X)⁻¹X′Y
- Partial coefficients measure each predictor’s effect controlling for others
- Adjusted R² penalizes added predictors; plain R² never decreases
- VIF diagnoses multicollinearity (flag > 5; severe > 10)
- Stepwise selection is exploratory — prefer AIC/BIC or Lasso
- Categorical variables need k−1 dummy indicators
12 / 12