STAT 101
M07 · L04
Module 7 · Lesson 4

Beyond Linear Regression

The straight line assumed a continuous, normal outcome. Real outcomes curve, or are yes/no, or are counts. This lesson generalises regression — all the way to the door of machine learning.

01 / 11
STAT 101
M07 · L04
Where the line breaks

One Model Wasn’t Enough

  • Curved relationships — a line underfits
  • Binary outcomes (yes/no) — a line predicts <0 and >1
  • Counts — non-negative, variance grows with the mean

Three fixes, then one framework that unifies them.

02 / 11
STAT 101
M07 · L04
Curves, still linear

Polynomial Regression

Add powers of x
y = \beta_0 + \beta_1 x + \dots + \beta_k x^k

Still linear in the coefficients, so OLS and every diagnostic apply. But high degrees overfit — degree 2 or 3 is usually plenty.

03 / 11
STAT 101
M07 · L04
Binary outcomes

Logistic Regression

The logit link
\ln\!\dfrac{p}{1-p} = \beta_0 + \beta_1 x_1 + \dots + \beta_p x_p
Reading it
The sigmoid keeps p in (0, 1). eβ is an odds ratio. Fit by maximum likelihood, not least squares.
04 / 11
STAT 101
M07 · L04
Count outcomes

Poisson for Counts

The log link
\ln \mathbb{E}[Y] = \mathbf{x}^\prime\boldsymbol{\beta}

The log link keeps the mean positive; eβ is a multiplicative rate factor. Variance > mean? Use negative binomial.

05 / 11
STAT 101
M07 · L04
The unifying idea

The GLM

Link function g
g\big(\mathbb{E}[Y]\big) = \mathbf{x}^\prime\boldsymbol{\beta}

Pick a distribution family and a link g. One framework, one fitting method (maximum likelihood).

06 / 11
STAT 101
M07 · L04
One family, three members

It’s All One Model

  • Linear — Gaussian family, identity link
  • Logistic — Binomial family, logit link
  • Poisson — Poisson family, log link

Ordinary regression is just the GLM whose link does nothing.

07 / 11
STAT 101
M07 · L04
Taming big models

Ridge & Lasso

Ridge — L2 penalty
\min\ \text{RSS} + \lambda \textstyle\sum_j \beta_j^2
Lasso — L1 penalty
\min\ \text{RSS} + \lambda \textstyle\sum_j |\beta_j|

Ridge shrinks smoothly; Lasso zeroes coefficients (feature selection); Elastic Net blends both. λ set by cross-validation.

08 / 11
STAT 101
M07 · L04
Where stats meets ML

The Bridge to ML

  • Logistic regression is a linear classifier
  • Regularization is the universal anti-overfitting tool
  • Choosing a loss + link is how modern models are specified
09 / 11
STAT 101
Beyond Linear Regression

Past the line

Four questions on GLMs, logistic and Poisson models, and regularization.

Question 1 of 0
Score 0/0

10 / 11
STAT 101
Summary
Recap

What you learned

  • Polynomial regression bends the fit but stays OLS — watch overfitting
  • Logistic models log-odds of a binary outcome; eβ is an odds ratio
  • Poisson models the log of a count; eβ is a rate factor
  • GLMs unify all three: a family + a link, g(E[Y]) = Xβ
  • Ridge/Lasso/Elastic Net trade bias for variance; λ by cross-validation
  • These are the bridge from regression to machine learning
11 / 11