ML 101
M02 · L03
Module 2

Taming Complexity

Your model fits the training data perfectly — but fails on new data. That’s overfitting. Regularization adds a penalty to keep weights small, trading a little training accuracy for much better generalization.

01 / 13
ML 101
M02 · L03
The Problem

The Overfitting Trap

A high-degree polynomial can pass through every training point — zero training error. But between those points it oscillates wildly, producing terrible predictions on unseen data.

Root Cause
When weights grow large, the model becomes hypersensitive to tiny input changes. Regularization keeps weights small.
02 / 13
ML 101
M02 · L03
The Idea

Penalize Big Weights

Add a penalty term to the cost function. Now the optimizer must balance fitting the data and keeping weights small. The parameter λ controls the trade-off.

Intuition
Small λ → nearly no penalty (risk overfitting). Large λ → strong penalty (risk underfitting).
03 / 13
ML 101
M02 · L03
L2 Penalty

Ridge Regression

L2 regularization adds the sum of squared weights to the cost. It shrinks all weights toward zero — smoothly and proportionally — but never exactly to zero.

Ridge Cost
J(\mathbf{w}) = \frac{1}{n}\sum_{i=1}^{n}(\hat{y}_i - y_i)^2 + \lambda \sum_{j=1}^{p} w_j^2
04 / 13
ML 101
M02 · L03
Closed Form

Ridge Solution

Like ordinary linear regression, Ridge has a closed-form solution. The λI term makes the matrix always invertible — even when features outnumber samples.

Ridge Normal Equation
\mathbf{w} = (\mathbf{X}^T \mathbf{X} + \lambda \mathbf{I})^{-1} \mathbf{X}^T \mathbf{y}
05 / 13
ML 101
M02 · L03
L1 Penalty

Lasso Regression

L1 regularization adds the sum of absolute values of weights. Its key superpower: it drives irrelevant weights exactly to zero — automatic feature selection.

Lasso Cost
J(\mathbf{w}) = \frac{1}{n}\sum_{i=1}^{n}(\hat{y}_i - y_i)^2 + \lambda \sum_{j=1}^{p} |w_j|
06 / 13
ML 101
M02 · L03
Geometry

Circle vs Diamond

L2’s constraint region is a circle (smooth, no corners). L1’s is a diamond (sharp corners on the axes). The loss contours hit diamond corners — where one weight is exactly zero.

L2 (Ridge)
Shrinks all
L1 (Lasso)
Zeros some
07 / 13
ML 101
M02 · L03
Best of Both

Elastic Net

Why choose? Elastic Net combines L1 and L2 penalties. It does feature selection like Lasso while maintaining Ridge’s stability when features are correlated.

Elastic Net
J(\mathbf{w}) = \text{MSE} + \lambda_1 \sum |w_j| + \lambda_2 \sum w_j^2
08 / 13
ML 101
M02 · L03
Tuning

Choosing λ

Too small → overfitting. Too large → underfitting. Use cross-validation: try a range of λ values, pick the one with lowest validation error.

Common Strategy
Search over λ = 0.001, 0.01, 0.1, 1, 10, 100 on a log scale. 5-fold CV is standard.
09 / 13
ML 101
M02 · L03
Deeper View

Bayesian Prior

Regularization is equivalent to placing a prior belief on weights. L2 = Gaussian prior (weights likely near zero). L1 = Laplace prior (weights likely exactly zero). The data updates this belief.

L2 Prior
Gaussian
L1 Prior
Laplace
10 / 13
ML 101
M02 · L03
In Practice

When to Use What

Start simple. If overfitting, add regularization:

  • Ridge — many features, all potentially useful
  • Lasso — suspect many features are irrelevant
  • Elastic Net — correlated features + sparsity needed
  • Always scale features before regularizing
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What you learned

Regularization fights overfitting by penalizing large weights. L2 (Ridge) shrinks all weights. L1 (Lasso) zeros irrelevant ones. Elastic Net combines both. λ controls the strength.

Next Module
13 / 13