Linear Regression
The foundation of machine learning. A straight line through data that predicts the future — simple, powerful, and the starting point for everything that follows.
Predicting with Data
Given past examples of inputs and outputs, can we find a pattern? Linear regression draws the best-fitting straight line through scattered data points.
The Linear Model
The simplest possible model — a straight line defined by just two numbers: a weight (slope) and a bias (intercept).
Multiple Features
Real problems have many inputs. Price depends on size, bedrooms, location, age… Linear regression handles them all with one weight per feature.
Mean Squared Error
How wrong is our line? MSE averages the squared distance between predictions and reality. Squaring penalizes big mistakes more than small ones.
The Loss Landscape
Plot MSE against every possible weight value and you get a bowl-shaped surface. The bottom of the bowl is the optimal solution. Our job: find it.
Gradient Descent
Start anywhere on the bowl. Compute the slope. Take a step downhill. Repeat. The learning rate α controls step size — too big and you overshoot, too small and you crawl.
Normal Equation
Why iterate when you can solve directly? The normal equation gives the exact optimal weights in one step — no learning rate, no iterations needed.
GD vs Normal Equation
Both find the same answer. The right choice depends on your data:
Feature Scaling
If one feature ranges 0–1 and another 0–1,000,000, gradient descent zigzags wildly. Normalize features to similar scales and it converges much faster.
Polynomial Regression
Data isn’t always linear. Add polynomial features like x² and x³ to fit curves — still using the same linear regression machinery under the hood.
What you learned
Linear regression fits a line to data by minimizing MSE. Gradient descent or normal equations find optimal weights. Feature scaling speeds convergence. Polynomial features handle curves.