Beyond Linear Regression
The straight line assumed a continuous, normal outcome. Real outcomes curve, or are yes/no, or are counts. This lesson generalises regression — all the way to the door of machine learning.
One Model Wasn’t Enough
- Curved relationships — a line underfits
- Binary outcomes (yes/no) — a line predicts <0 and >1
- Counts — non-negative, variance grows with the mean
Three fixes, then one framework that unifies them.
Polynomial Regression
Still linear in the coefficients, so OLS and every diagnostic apply. But high degrees overfit — degree 2 or 3 is usually plenty.
Logistic Regression
Poisson for Counts
The log link keeps the mean positive; eβ is a multiplicative rate factor. Variance > mean? Use negative binomial.
The GLM
Pick a distribution family and a link g. One framework, one fitting method (maximum likelihood).
It’s All One Model
- Linear — Gaussian family, identity link
- Logistic — Binomial family, logit link
- Poisson — Poisson family, log link
Ordinary regression is just the GLM whose link does nothing.
Ridge & Lasso
Ridge shrinks smoothly; Lasso zeroes coefficients (feature selection); Elastic Net blends both. λ set by cross-validation.
The Bridge to ML
- Logistic regression is a linear classifier
- Regularization is the universal anti-overfitting tool
- Choosing a loss + link is how modern models are specified
What you learned
- Polynomial regression bends the fit but stays OLS — watch overfitting
- Logistic models log-odds of a binary outcome; eβ is an odds ratio
- Poisson models the log of a count; eβ is a rate factor
- GLMs unify all three: a family + a link, g(E[Y]) = Xβ
- Ridge/Lasso/Elastic Net trade bias for variance; λ by cross-validation
- These are the bridge from regression to machine learning