ML 101
M02 · L01
Module 2

Linear Regression

The foundation of machine learning. A straight line through data that predicts the future — simple, powerful, and the starting point for everything that follows.

01 / 13
ML 101
M02 · L01
The Setup

Predicting with Data

Given past examples of inputs and outputs, can we find a pattern? Linear regression draws the best-fitting straight line through scattered data points.

Example
Given house sizes and prices, predict the price of a new house.
02 / 13
ML 101
M02 · L01
The Equation

The Linear Model

The simplest possible model — a straight line defined by just two numbers: a weight (slope) and a bias (intercept).

Linear Model
\hat{y} = w_1 x + w_0
03 / 13
ML 101
M02 · L01
Scaling Up

Multiple Features

Real problems have many inputs. Price depends on size, bedrooms, location, age… Linear regression handles them all with one weight per feature.

Multiple Regression
\hat{y} = w_0 + w_1 x_1 + w_2 x_2 + \cdots + w_n x_n
04 / 13
ML 101
M02 · L01
Measuring Error

Mean Squared Error

How wrong is our line? MSE averages the squared distance between predictions and reality. Squaring penalizes big mistakes more than small ones.

Cost Function
J(\mathbf{w}) = \frac{1}{n}\sum_{i=1}^{n}(\hat{y}_i - y_i)^2
05 / 13
ML 101
M02 · L01
Visualization

The Loss Landscape

Plot MSE against every possible weight value and you get a bowl-shaped surface. The bottom of the bowl is the optimal solution. Our job: find it.

Key Insight
For linear regression, this bowl is always convex — there’s exactly one global minimum. No traps, no dead ends.
06 / 13
ML 101
M02 · L01
Optimization

Gradient Descent

Start anywhere on the bowl. Compute the slope. Take a step downhill. Repeat. The learning rate α controls step size — too big and you overshoot, too small and you crawl.

Update Rule
w \leftarrow w - \alpha \frac{\partial J}{\partial w}
07 / 13
ML 101
M02 · L01
The Shortcut

Normal Equation

Why iterate when you can solve directly? The normal equation gives the exact optimal weights in one step — no learning rate, no iterations needed.

Closed-Form Solution
\mathbf{w} = (\mathbf{X}^T \mathbf{X})^{-1} \mathbf{X}^T \mathbf{y}
08 / 13
ML 101
M02 · L01
Trade-offs

GD vs Normal Equation

Both find the same answer. The right choice depends on your data:

Normal Equation
n < 10,000
Gradient Descent
n > 10,000
09 / 13
ML 101
M02 · L01
Practical Tip

Feature Scaling

If one feature ranges 0–1 and another 0–1,000,000, gradient descent zigzags wildly. Normalize features to similar scales and it converges much faster.

Before Scaling
1000+ iters
After Scaling
~50 iters
10 / 13
ML 101
M02 · L01
Beyond Lines

Polynomial Regression

Data isn’t always linear. Add polynomial features like x² and x³ to fit curves — still using the same linear regression machinery under the hood.

Polynomial Model
\hat{y} = w_0 + w_1 x + w_2 x^2 + w_3 x^3
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What you learned

Linear regression fits a line to data by minimizing MSE. Gradient descent or normal equations find optimal weights. Feature scaling speeds convergence. Polynomial features handle curves.

13 / 13