LA 101
M11 · L01
Module 11: Machine Learning Applications

Linear Models

From regression to support vector machines, the most powerful ideas in machine learning are expressions of linear algebra. The design matrix X is the backbone of every linear model — organize data into it, and the rest is matrix arithmetic.

01 / 12
LA 101
M11 · L01
Linear Regression

The Design Matrix

Given n examples with p features each, arrange them into an n×p design matrix X. The model predicts outputs as y = Xβ + ε. Finding β means solving a least-squares problem — minimize the squared distance from y to the column space of X.

Linear Model
\mathbf{y}=X\boldsymbol{\beta}+\boldsymbol{\varepsilon}
02 / 12
LA 101
M11 · L01
Geometric View

Regression is Projection

The OLS solution is the orthogonal projection of y onto col(X). Setting the gradient of ‖y−Xβ‖² to zero gives the normal equations XᵀXβ = Xᵀy, with the unique solution:

OLS Solution
\hat{\boldsymbol{\beta}}=(X^T X)^{-1}X^T\mathbf{y}

The residual r = y − ŷ is perpendicular to every column of X: Xᵀr = 0.

03 / 12
LA 101
M11 · L01
Regularization: L2

Ridge Regression

When XᵀX is ill-conditioned, tiny changes in y cause huge swings in β̂. Ridge adds λI to guarantee invertibility and shrink coefficients toward zero:

Ridge Solution
\hat{\boldsymbol{\beta}}_{\text{ridge}}=(X^T X+\lambda I)^{-1}X^T\mathbf{y}

Via SVD: each component σₖ is shrunk by σₖ²/(σₖ²+λ). Larger λ means more shrinkage.

04 / 12
LA 101
M11 · L01
Regularization: L1

Lasso and Sparsity

Replace the L2 penalty with L1: ‖β‖₁ = Σ|βⱼ|. The L1 ball is a polytope — its extreme points lie on coordinate axes. The optimum often falls at a corner, setting many coefficients exactly to zero. Lasso performs automatic variable selection.

‖β‖₂²
Ridge Ball (smooth)
‖β‖₁
Lasso Ball (corners)
05 / 12
LA 101
M11 · L01
Ridge vs. Lasso

Choosing Your Regularizer

  • Ridge: closed-form, all coefs shrink but none zero, good when all features matter
  • Lasso: no closed-form, many coefs exactly zero, good for feature selection
  • Elastic Net: combines both — sparse and stable
  • Both are convex — guaranteed global minimum
06 / 12
LA 101
M11 · L01
The Kernel Trick

Inner Products Without Mapping

Map x → φ(x) to a richer feature space, but never compute φ explicitly. Use only the kernel k(xᵢ,xⱼ) = φ(xᵢ)ᵀφ(xⱼ). Any algorithm depending only on inner products can be kernelized — replacing XᵀX with the kernel matrix K.

Common Kernels
RBF: k(x,z)=exp(−‖x−z‖²/2σ²) · Polynomial: (xᵀz+c)^d · Linear: xᵀz
07 / 12
LA 101
M11 · L01
Support Vector Machines

Maximum Margin Hyperplane

An SVM finds the separating hyperplane wᵀx + b = 0 with the largest margin 2/‖w‖. Maximizing margin = minimizing ‖w‖². This is a quadratic program:

Hard-Margin SVM
\min_{\mathbf{w},b}\;\tfrac{1}{2}\|\mathbf{w}\|^2\quad\text{s.t.}\quad y_i(\mathbf{w}^T\mathbf{x}_i+b)\geq 1
08 / 12
LA 101
M11 · L01
SVM Dual

Support Vectors and the Dual

The Lagrangian dual of the SVM QP depends only on inner products xᵢᵀxⱼ — making kernelization immediate. The optimal w = Σᵢ αᵢ yᵢ xᵢ is a sparse combination: only the support vectors (training points at the margin boundary, where αᵢ > 0) contribute.

QP
Primal Form
αᵢxᵀx
Dual (kernelizable)
09 / 12
LA 101
Linear models

Check what stuck

Four questions on the linear algebra behind regression, regularization and SVMs.

Question 1 of 0
Score 0/0

10 / 12
LA 101
Key Takeaways
Summary

Key Takeaways

  • Linear regression = orthogonal projection: β̂ = (XᵀX)⁻¹Xᵀy
  • Ridge (L2 penalty): stabilizes ill-conditioned XᵀX, shrinks via σ²/(σ²+λ)
  • Lasso (L1 penalty): polytope geometry produces exact sparsity — automatic feature selection
  • Kernel trick: replace inner products with k(xᵢ,xⱼ) — nonlinear model without explicit mapping
  • SVM: maximum-margin QP whose dual depends only on inner products — naturally kernelizable
11 / 12
LA 101
Up Next
Coming Up

M11-L2: Matrix Factorization

We've seen how linear algebra powers regression and classification. Next we'll explore how matrix factorization — NMF, SVD-based collaborative filtering, and latent semantic analysis — learns structure hidden in large data matrices.

12 / 12