Least Squares
When a system is overdetermined — more equations than unknowns — we can't solve it exactly. Instead we find the best approximation by minimizing the squared error.
More Equations Than Unknowns
A is m×n with m > n. The vector b usually lies outside the column space of A. No x satisfies Ax = b exactly.
Project onto C(A)
The closest point in C(A) to b is its orthogonal projection p = Ax̂. The residual r = b − p is perpendicular to every column of A.
Normal Equations
Setting Aᵀ(b − Ax̂) = 0 gives the normal equations. When A has full column rank, the unique solution is:
P = A(AᵀA)⁻¹Aᵀ
Linear Regression
Fitting b = x₁ + x₂t to m data points is exactly least squares. The design matrix A has a column of ones and a column of t values.
- Normal equations → intercept x̂₁ and slope x̂₂
- Extends to polynomials, multiple predictors
- Weighted regression: minimize ‖W(b − Ax)‖²
- Residual ‖b − Ax̂‖² / (m−n) estimates noise variance
QR beats AᵀA
Forming AᵀA squares the condition number. Use QR instead: write A = QR, then solve Rx̂ = Qᵀb by back-substitution.
Least Squares in Practice
- Channel estimation — pilot symbols → ĥ = (ΦᵀΦ)⁻¹Φᵀy
- Beamforming — weights minimizing ‖d − Aw‖²
- System identification — FIR/AR model fitting from I/O data
- GPS positioning — overdetermined range equations
- SVD — minimum-norm solution for rank-deficient A
Module 5 Done!
You've completed Module 5: Orthogonality — projections, Gram-Schmidt, orthogonal matrices, and least squares. These tools form the backbone of modern data analysis and signal processing.