Reading
Stories Mode

Regularization

~14 min read Lesson 3 of Module 2

The Overfitting Problem

In the previous lesson on linear regression, we saw that polynomial features let us fit curves. But there’s a dangerous trap: a high-degree polynomial can pass through every single training point — achieving zero training error — while producing wildly wrong predictions on new data. The model has memorized the noise instead of learning the pattern.

The root cause? Large weights. When weights grow unconstrained, the model becomes hypersensitive — tiny changes in input cause enormous swings in output. Regularization solves this by adding a penalty that discourages large weights.

The Regularized Cost Function

The idea is elegant: modify the cost function so the optimizer must balance two competing objectives. Fit the data well (minimize MSE) and keep weights small (minimize the penalty). The regularization parameter λ controls the trade-off — higher λ means stronger regularization.

L2 Regularization: Ridge Regression

Ridge regression adds the sum of squared weights to the cost function. This penalty grows quadratically as any weight increases, creating strong pressure to keep all weights small and evenly distributed.

Ridge Cost Function
J(\mathbf{w}) = \frac{1}{n}\sum_{i=1}^{n}(\hat{y}_i - y_i)^2 + \lambda \sum_{j=1}^{p} w_j^2
MSE plus λ times the sum of squared weights. The penalty shrinks weights toward zero but never exactly to zero.

Ridge has a clean closed-form solution that extends the normal equation. The λI term ensures the matrix is always invertible — even when you have more features than data points.

Ridge Solution
\mathbf{w} = (\mathbf{X}^T \mathbf{X} + \lambda \mathbf{I})^{-1} \mathbf{X}^T \mathbf{y}
The λI term regularizes the matrix inversion, guaranteeing a unique solution even in ill-conditioned problems.

L1 Regularization: Lasso Regression

Lasso regression (Least Absolute Shrinkage and Selection Operator) adds the sum of absolute values of weights. The critical difference from Ridge: L1 can drive weights exactly to zero, effectively removing irrelevant features from the model.

Lasso Cost Function
J(\mathbf{w}) = \frac{1}{n}\sum_{i=1}^{n}(\hat{y}_i - y_i)^2 + \lambda \sum_{j=1}^{p} |w_j|
The absolute value penalty creates sparse solutions — many weights become exactly zero, performing automatic feature selection.
Why Does L1 Zero Out Weights?

Geometry explains it. L2’s constraint region is a smooth circle — the loss contours typically touch it at non-axis points (all weights nonzero). L1’s constraint region is a diamond with sharp corners on the axes — the loss contours are much more likely to touch at a corner where one or more weights are exactly zero.

Elastic Net: Combining L1 and L2

What if you want feature selection (L1) and the stability of Ridge (L2)? Elastic Net combines both penalties. This is especially useful when features are correlated — Lasso tends to arbitrarily pick one from a group of correlated features, while Elastic Net keeps them together.

Elastic Net
J(\mathbf{w}) = \text{MSE} + \lambda_1 \sum |w_j| + \lambda_2 \sum w_j^2
Combines L1 (sparsity) and L2 (stability) penalties. Two hyperparameters: λ₁ controls sparsity, λ₂ controls shrinkage.

Choosing λ

The regularization strength λ is a hyperparameter — it’s not learned from data; you must choose it. Too small and overfitting persists. Too large and the model underfits, ignoring the data in favor of near-zero weights.

The standard approach is cross-validation: try a range of λ values (typically on a logarithmic scale: 0.001, 0.01, 0.1, 1, 10, 100), evaluate each with k-fold cross-validation, and select the one with lowest validation error.

The Bayesian Perspective

Regularization has a beautiful probabilistic interpretation. Adding a penalty is mathematically equivalent to placing a prior distribution on the weights before seeing data. L2 regularization corresponds to a Gaussian prior (we believe weights are normally distributed around zero). L1 corresponds to a Laplace prior (sharply peaked at zero, making exact zeros likely). The data then updates these beliefs.

Practical Guidelines

Use Ridge when you believe all features contribute and want stable predictions. Use Lasso when you suspect many features are irrelevant and want automatic feature selection. Use Elastic Net when features are correlated. Always scale features before applying regularization — otherwise the penalty unfairly penalizes features on larger scales.

With linear models and regularization under your belt, you’re ready for a fundamentally different approach — decision trees, which learn rules instead of equations, and ensemble methods that combine many weak learners into powerful predictors.

Key Takeaways
  • Regularization fights overfitting by adding a penalty for large weights to the cost function, controlled by λ.
  • L2 (Ridge) adds squared weights — shrinks all weights toward zero but never exactly to zero. Has a closed-form solution.
  • L1 (Lasso) adds absolute weights — drives irrelevant weights exactly to zero, performing automatic feature selection.
  • Elastic Net combines L1 and L2, offering both sparsity and stability for correlated features.
  • Choose λ via cross-validation. Regularization is equivalent to placing a Bayesian prior on weights.
PreviousLogistic Regression Module Overview Next ModuleDecision Trees