Reading
Stories Mode

Beyond Linear Regression

~30 min read Lesson 4 of 4 in Module 7

One Model Was Never Enough

The three lessons before this one built, diagnosed, and repaired the linear model Y = β₀ + β₁X₁ + … + ε. It is the workhorse of statistics, but it makes two assumptions that the real world routinely breaks: that the relationship is a straight line, and that the response is a continuous quantity with roughly normal, constant-variance errors. Ask it to predict whether a patient survives, how many calls arrive per hour, or a curve that bends, and it strains or fails outright.

This lesson generalises the linear model in two directions. First, we keep least squares but let the predictors curve. Then we keep the linear predictor but change what it is connected to — a probability, a count, a rate — through the generalized linear model. Finally we add a penalty that keeps large models honest, the idea that carries regression directly into machine learning.

Polynomial Regression: Curves, Still Linear

The quickest way past the straight-line assumption is to add powers of a predictor as new columns — x², x³, and so on — and fit them with ordinary least squares:

Polynomial Regression
y = \beta_0 + \beta_1 x + \beta_2 x^2 + \dots + \beta_k x^k + \varepsilon
A degree-k polynomial. Crucially it is still linear in the coefficients β, so OLS, R², and every diagnostic from the last lesson apply unchanged — only the design matrix has extra columns.

The catch is overfitting: a high-degree polynomial will chase every wiggle in the sample and swing wildly between and beyond the data points, especially at the edges. In practice degree 2 or 3 is almost always enough; anything higher is usually a sign you want a genuinely non-linear method (splines, covered in Module 8) rather than a taller polynomial.

Logistic Regression: Modelling a Probability

When the outcome is binary — survived or not, clicked or not, defaulted or not — fitting a straight line to the 0/1 response is nonsense: it predicts probabilities above 1 and below 0, and its errors are neither normal nor constant-variance. Logistic regression fixes this by modelling not the outcome but the log-odds of the outcome as a linear function:

The Logit Link
\operatorname{logit}(p) = \ln\!\left(\dfrac{p}{1-p}\right) = \beta_0 + \beta_1 x_1 + \dots + \beta_p x_p
The logit — the natural log of the odds p/(1−p) — is modelled as linear. Inverting it gives the S-shaped logistic (sigmoid) curve p = 1/(1 + e−x′β), which is bounded to (0, 1) as a probability must be.

Because the model is linear in the log-odds, each coefficient has a clean reading: eβj is the odds ratio for a one-unit change in xj. Coefficients are fit by maximum likelihood rather than least squares, and significance is judged with z- (Wald) tests and likelihood-ratio tests instead of the F-test. Logistic regression is simultaneously the workhorse of epidemiology and credit scoring and the simplest classifier in machine learning — the same equation, read from two traditions.

Poisson Regression: Modelling Counts

Counts — accidents per intersection, emails per hour, defects per batch — are non-negative integers whose variance grows with their mean, so they too break the linear model. Poisson regression models the log of the expected count as linear:

The Log Link
\ln\big(\mathbb{E}[Y \mid \mathbf{x}]\big) = \mathbf{x}^\prime\boldsymbol{\beta}, \qquad Y \sim \text{Poisson}(\lambda)
The log link keeps the fitted mean positive and turns multiplicative effects into additive coefficients: eβj is the factor by which the expected count multiplies per unit of xj. When the variance exceeds the mean (overdispersion), a negative-binomial model is the usual next step.

The Unifying Idea: Generalized Linear Models

Polynomial, logistic, and Poisson regression are not three unrelated tricks. They are three instances of one framework, the generalized linear model (GLM), which keeps the familiar linear predictor β₀ + β₁x₁ + … and connects it to the mean of the response through a link function g:

The GLM
g\big(\mathbb{E}[Y]\big) = \beta_0 + \beta_1 x_1 + \dots + \beta_p x_p
A GLM is chosen by naming two things: the distribution family of the response and the link function g that maps its mean onto the linear predictor.
One Framework, Three Familiar Members

Linear regression — Normal (Gaussian) family, identity link g(μ) = μ. The ordinary model is just the GLM whose link does nothing.

Logistic regression — Binomial family, logit link g(μ) = ln(μ/(1−μ)). For binary outcomes.

Poisson regression — Poisson family, log link g(μ) = lnμ. For counts.

Seeing them as one family is the point: choose the distribution that matches your data and the link that keeps predictions in range, and the same fitting machinery (maximum likelihood) and the same style of inference carry over.

Regularization: Ridge, Lasso, and Elastic Net

With many predictors — especially when they are correlated or nearly as numerous as the observations — least squares overfits and its coefficients become large and unstable. Regularization adds a penalty on coefficient size to the fitting criterion, deliberately accepting a little bias to buy a large reduction in variance.

Ridge Regression (L2 penalty)
\hat{\boldsymbol{\beta}}_{\text{ridge}} = \arg\min_{\boldsymbol{\beta}} \sum_{i=1}^{n}(y_i - \mathbf{x}_i^\prime\boldsymbol{\beta})^2 + \lambda \sum_{j=1}^{p} \beta_j^2
Ridge adds λ times the sum of squared coefficients. It shrinks all coefficients smoothly toward zero — never exactly to zero — and is especially good at taming multicollinearity.
Lasso Regression (L1 penalty)
\hat{\boldsymbol{\beta}}_{\text{lasso}} = \arg\min_{\boldsymbol{\beta}} \sum_{i=1}^{n}(y_i - \mathbf{x}_i^\prime\boldsymbol{\beta})^2 + \lambda \sum_{j=1}^{p} |\beta_j|
Lasso penalises the sum of absolute coefficients. Its geometry drives some coefficients exactly to zero, so it performs automatic feature selection — the model chooses which predictors to keep.

Elastic Net mixes the two penalties, λ₁Σ|βj| + λ₂Σβj², inheriting Lasso’s sparsity while handling groups of correlated predictors more gracefully than Lasso alone. In all three, the strength λ is a tuning knob chosen by cross-validation: λ = 0 recovers ordinary least squares, and larger λ means a simpler, more heavily shrunk model.

The Bridge to Machine Learning

This lesson is where classical regression and machine learning meet. Logistic regression is a linear classifier, and it is the last layer of a neural network. Regularization — the L1 and L2 penalties above — is not a statistical footnote but the default defence against overfitting across all of modern ML, from Lasso to weight decay in deep networks. And the GLM habit of choosing a loss that matches the data and a link that keeps predictions in range is exactly how a modern model is specified. The line from “fit a straight line” to “train a classifier” is shorter, and straighter, than it looks.

Key Takeaways
  • Polynomial regression bends the fit by adding x², x³, … as predictors; it stays linear in the coefficients, so OLS still applies — but high degrees overfit badly.
  • Logistic regression models the log-odds of a binary outcome as linear; the sigmoid keeps predictions in (0, 1) and eβ is an odds ratio. It is fit by maximum likelihood.
  • Poisson regression models the log of an expected count as linear; eβ is a multiplicative rate factor. Overdispersion points to the negative-binomial model.
  • Generalized linear models unify all of these: pick a distribution family and a link function g with g(E[Y]) = Xβ. Linear regression is the identity-link Gaussian special case.
  • Regularization trades bias for variance: Ridge (L2) shrinks smoothly, Lasso (L1) zeroes coefficients for feature selection, Elastic Net blends both. Strength λ is set by cross-validation.
  • These are the bridge to machine learning: logistic regression is a classifier, regularization is the universal anti-overfitting tool, and the GLM mindset is how modern models are specified.
Previous Regression Diagnostics Module Overview Next Why Non-Parametric?