The Overfitting Problem
In the previous lesson on linear regression, we saw that polynomial features let us fit curves. But there’s a dangerous trap: a high-degree polynomial can pass through every single training point — achieving zero training error — while producing wildly wrong predictions on new data. The model has memorized the noise instead of learning the pattern.
The root cause? Large weights. When weights grow unconstrained, the model becomes hypersensitive — tiny changes in input cause enormous swings in output. Regularization solves this by adding a penalty that discourages large weights.
The Regularized Cost Function
The idea is elegant: modify the cost function so the optimizer must balance two competing objectives. Fit the data well (minimize MSE) and keep weights small (minimize the penalty). The regularization parameter λ controls the trade-off — higher λ means stronger regularization.
L2 Regularization: Ridge Regression
Ridge regression adds the sum of squared weights to the cost function. This penalty grows quadratically as any weight increases, creating strong pressure to keep all weights small and evenly distributed.
Ridge has a clean closed-form solution that extends the normal equation. The λI term ensures the matrix is always invertible — even when you have more features than data points.
L1 Regularization: Lasso Regression
Lasso regression (Least Absolute Shrinkage and Selection Operator) adds the sum of absolute values of weights. The critical difference from Ridge: L1 can drive weights exactly to zero, effectively removing irrelevant features from the model.
Geometry explains it. L2’s constraint region is a smooth circle — the loss contours typically touch it at non-axis points (all weights nonzero). L1’s constraint region is a diamond with sharp corners on the axes — the loss contours are much more likely to touch at a corner where one or more weights are exactly zero.
Elastic Net: Combining L1 and L2
What if you want feature selection (L1) and the stability of Ridge (L2)? Elastic Net combines both penalties. This is especially useful when features are correlated — Lasso tends to arbitrarily pick one from a group of correlated features, while Elastic Net keeps them together.
Choosing λ
The regularization strength λ is a hyperparameter — it’s not learned from data; you must choose it. Too small and overfitting persists. Too large and the model underfits, ignoring the data in favor of near-zero weights.
The standard approach is cross-validation: try a range of λ values (typically on a logarithmic scale: 0.001, 0.01, 0.1, 1, 10, 100), evaluate each with k-fold cross-validation, and select the one with lowest validation error.
The Bayesian Perspective
Regularization has a beautiful probabilistic interpretation. Adding a penalty is mathematically equivalent to placing a prior distribution on the weights before seeing data. L2 regularization corresponds to a Gaussian prior (we believe weights are normally distributed around zero). L1 corresponds to a Laplace prior (sharply peaked at zero, making exact zeros likely). The data then updates these beliefs.
Use Ridge when you believe all features contribute and want stable predictions. Use Lasso when you suspect many features are irrelevant and want automatic feature selection. Use Elastic Net when features are correlated. Always scale features before applying regularization — otherwise the penalty unfairly penalizes features on larger scales.
- Regularization fights overfitting by adding a penalty for large weights to the cost function, controlled by λ.
- L2 (Ridge) adds squared weights — shrinks all weights toward zero but never exactly to zero. Has a closed-form solution.
- L1 (Lasso) adds absolute weights — drives irrelevant weights exactly to zero, performing automatic feature selection.
- Elastic Net combines L1 and L2, offering both sparsity and stability for correlated features.
- Choose λ via cross-validation. Regularization is equivalent to placing a Bayesian prior on weights.