Taming Complexity
Your model fits the training data perfectly — but fails on new data. That’s overfitting. Regularization adds a penalty to keep weights small, trading a little training accuracy for much better generalization.
The Overfitting Trap
A high-degree polynomial can pass through every training point — zero training error. But between those points it oscillates wildly, producing terrible predictions on unseen data.
Penalize Big Weights
Add a penalty term to the cost function. Now the optimizer must balance fitting the data and keeping weights small. The parameter λ controls the trade-off.
Ridge Regression
L2 regularization adds the sum of squared weights to the cost. It shrinks all weights toward zero — smoothly and proportionally — but never exactly to zero.
Ridge Solution
Like ordinary linear regression, Ridge has a closed-form solution. The λI term makes the matrix always invertible — even when features outnumber samples.
Lasso Regression
L1 regularization adds the sum of absolute values of weights. Its key superpower: it drives irrelevant weights exactly to zero — automatic feature selection.
Circle vs Diamond
L2’s constraint region is a circle (smooth, no corners). L1’s is a diamond (sharp corners on the axes). The loss contours hit diamond corners — where one weight is exactly zero.
Elastic Net
Why choose? Elastic Net combines L1 and L2 penalties. It does feature selection like Lasso while maintaining Ridge’s stability when features are correlated.
Choosing λ
Too small → overfitting. Too large → underfitting. Use cross-validation: try a range of λ values, pick the one with lowest validation error.
Bayesian Prior
Regularization is equivalent to placing a prior belief on weights. L2 = Gaussian prior (weights likely near zero). L1 = Laplace prior (weights likely exactly zero). The data updates this belief.
When to Use What
Start simple. If overfitting, add regularization:
- Ridge — many features, all potentially useful
- Lasso — suspect many features are irrelevant
- Elastic Net — correlated features + sparsity needed
- Always scale features before regularizing
What you learned
Regularization fights overfitting by penalizing large weights. L2 (Ridge) shrinks all weights. L1 (Lasso) zeros irrelevant ones. Elastic Net combines both. λ controls the strength.