Linear Models
From regression to support vector machines, the most powerful ideas in machine learning are expressions of linear algebra. The design matrix X is the backbone of every linear model — organize data into it, and the rest is matrix arithmetic.
The Design Matrix
Given n examples with p features each, arrange them into an n×p design matrix X. The model predicts outputs as y = Xβ + ε. Finding β means solving a least-squares problem — minimize the squared distance from y to the column space of X.
Regression is Projection
The OLS solution is the orthogonal projection of y onto col(X). Setting the gradient of ‖y−Xβ‖² to zero gives the normal equations XᵀXβ = Xᵀy, with the unique solution:
The residual r = y − ŷ is perpendicular to every column of X: Xᵀr = 0.
Ridge Regression
When XᵀX is ill-conditioned, tiny changes in y cause huge swings in β̂. Ridge adds λI to guarantee invertibility and shrink coefficients toward zero:
Via SVD: each component σₖ is shrunk by σₖ²/(σₖ²+λ). Larger λ means more shrinkage.
Lasso and Sparsity
Replace the L2 penalty with L1: ‖β‖₁ = Σ|βⱼ|. The L1 ball is a polytope — its extreme points lie on coordinate axes. The optimum often falls at a corner, setting many coefficients exactly to zero. Lasso performs automatic variable selection.
Choosing Your Regularizer
- Ridge: closed-form, all coefs shrink but none zero, good when all features matter
- Lasso: no closed-form, many coefs exactly zero, good for feature selection
- Elastic Net: combines both — sparse and stable
- Both are convex — guaranteed global minimum
Inner Products Without Mapping
Map x → φ(x) to a richer feature space, but never compute φ explicitly. Use only the kernel k(xᵢ,xⱼ) = φ(xᵢ)ᵀφ(xⱼ). Any algorithm depending only on inner products can be kernelized — replacing XᵀX with the kernel matrix K.
Maximum Margin Hyperplane
An SVM finds the separating hyperplane wᵀx + b = 0 with the largest margin 2/‖w‖. Maximizing margin = minimizing ‖w‖². This is a quadratic program:
Support Vectors and the Dual
The Lagrangian dual of the SVM QP depends only on inner products xᵢᵀxⱼ — making kernelization immediate. The optimal w = Σᵢ αᵢ yᵢ xᵢ is a sparse combination: only the support vectors (training points at the margin boundary, where αᵢ > 0) contribute.
Key Takeaways
- Linear regression = orthogonal projection: β̂ = (XᵀX)⁻¹Xᵀy
- Ridge (L2 penalty): stabilizes ill-conditioned XᵀX, shrinks via σ²/(σ²+λ)
- Lasso (L1 penalty): polytope geometry produces exact sparsity — automatic feature selection
- Kernel trick: replace inner products with k(xᵢ,xⱼ) — nonlinear model without explicit mapping
- SVM: maximum-margin QP whose dual depends only on inner products — naturally kernelizable
M11-L2: Matrix Factorization
We've seen how linear algebra powers regression and classification. Next we'll explore how matrix factorization — NMF, SVD-based collaborative filtering, and latent semantic analysis — learns structure hidden in large data matrices.