The Foundation of Machine Learning
Linear regression is the simplest and most widely used machine learning model. It finds the best-fitting straight line through data points — given a set of inputs, it predicts a continuous output. Despite its simplicity, the core concepts you’ll learn here — cost functions, gradient descent, optimization — apply directly to neural networks with millions of parameters.
If you understand linear regression deeply, you understand the engine that drives all of supervised learning.
The Linear Model
At its core, linear regression is a straight line: given an input x, predict an output ŷ. The model has two learnable parameters — a weight (the slope, controlling how much x affects the prediction) and a bias (the y-intercept, the prediction when x is zero).
Multiple Linear Regression
Real-world problems rarely depend on a single input. House prices depend on size, number of bedrooms, location, age, and more. Multiple linear regression extends the model to handle any number of features — each gets its own weight, and the model learns how much each feature matters.
In matrix notation, multiple regression becomes ŷ = Xw, where X is the data matrix and w is the weight vector. This compact form makes implementation efficient and is how libraries like NumPy compute it in practice.
The Cost Function: Mean Squared Error
How do we know if our line is any good? We need a way to measure how wrong our predictions are. The most common measure is Mean Squared Error (MSE) — it computes the average of the squared differences between predicted and actual values.
Why squaring? It makes all errors positive (no cancellation between over- and under-predictions), it’s differentiable (needed for optimization), and it penalizes large errors more than small ones.
The Loss Landscape
If you plot MSE as a function of the weight values, you get a bowl-shaped surface (mathematically, a convex paraboloid). The bottom of the bowl is the set of weights that minimize the error — the optimal solution. For linear regression, this bowl has exactly one global minimum. No local minima traps, no saddle points — just a clean, smooth descent to the best answer.
Finding the Best Weights: Gradient Descent
Gradient descent is the workhorse of machine learning optimization. Imagine standing on the side of that bowl in fog. You can’t see the bottom, but you can feel which direction slopes downward. Take a step in that direction. Feel the slope again. Step again. Repeat until you reach the bottom.
The learning rate α controls step size. Too large and you overshoot the minimum, bouncing back and forth or even diverging. Too small and training takes forever, creeping toward the solution at a glacial pace. Finding the right learning rate is one of the key practical skills in ML.
The Normal Equation: A Direct Solution
For linear regression specifically, there’s a shortcut. Instead of iterating, we can compute the exact optimal weights in one step using calculus — set the gradient to zero and solve.
Use the normal equation when you have fewer than ~10,000 features — it’s exact and requires no hyperparameter tuning. Use gradient descent for larger feature sets — matrix inversion is O(n³), which becomes prohibitively expensive for very high-dimensional data.
Feature Scaling
When features live on wildly different scales — say, square footage (0–5,000) and number of bedrooms (0–10) — the loss landscape becomes an elongated ellipse instead of a nice round bowl. Gradient descent zigzags inefficiently along the narrow direction.
The fix is simple: normalize features to similar ranges. Common approaches include min-max scaling (map to [0,1]) and standardization (subtract mean, divide by standard deviation). After scaling, gradient descent converges dramatically faster — often 20× fewer iterations.
Polynomial Regression
What if the relationship isn’t linear? A straight line can’t capture a U-shaped or exponential pattern. The trick: create polynomial features — add x², x³, etc. — and apply linear regression to the expanded feature set.
Be careful: high-degree polynomials fit the training data perfectly but generalize poorly to new data. This is overfitting — and it sets the stage for our next topic: regularization, which controls model complexity to prevent it.
- Linear regression fits a line (or hyperplane) by minimizing Mean Squared Error — the average of squared prediction errors.
- Gradient descent iteratively adjusts weights by following the steepest downhill direction of the loss surface.
- The normal equation gives exact optimal weights directly — but doesn’t scale to very large feature sets due to O(n³) matrix inversion.
- Feature scaling normalizes inputs to similar ranges, dramatically speeding up gradient descent convergence.
- Polynomial regression fits curves by adding polynomial features to the same linear framework — but watch out for overfitting.