LA 101
M07 · L03
Module 7: Matrix Decompositions
Singular Value Decomposition
A = UΣVᵀ for any matrix — any shape, any rank. SVD reveals hidden geometry, powers image compression, and is the backbone of modern data science.
01 / 11
LA 101
M07 · L03
The Three Factors
U, Σ, V — Each Has a Role
U
m × m orthogonal, left singular vectors
Σ
diagonal, singular values σ₁ ≥ σ₂ ≥ 0
V
n × n orthogonal, right singular vectors
Universal
Works for any m × n matrix — square, rectangular, singular, rank-deficient. LU requires square; QR needs full column rank. SVD needs nothing.
02 / 11
LA 101
M07 · L03
SVD Formula
The Full Decomposition
Singular Value Decomposition
A = U\Sigma V^T,\quad \sigma_1 \geq \sigma_2 \geq \cdots \geq \sigma_r > 0
Outer product form
A = σ₁u₁v₁ᵀ + σ₂u₂v₂ᵀ + ⋯ + σᵣuᵣvᵣᵀ — a sum of r rank-one "building blocks," each weighted by a singular value.
03 / 11
LA 101
M07 · L03
Geometry
Rotate → Stretch → Rotate
Every matrix A acts as three steps: Vᵀ rotates the input, Σ stretches along coordinate axes by the singular values, U rotates the output.
- Vᵀ: rigid rotation — preserves all lengths and angles
- Σ: stretches axis i by σᵢ, collapses axes with σᵢ = 0
- U: rigid rotation in the output space
- Singular values tell you which directions A amplifies most
04 / 11
LA 101
M07 · L03
Four Fundamental Subspaces
SVD Reveals Everything
- Column space: first r left singular vectors u₁, …, uᵣ
- Left null space: remaining uᵣ₊₁, …, uₘ (zero singular values)
- Row space: first r right singular vectors v₁, …, vᵣ
- Null space: remaining vᵣ₊₁, …, vₙ
- Rank: number of non-zero singular values
- Condition number: κ₂(A) = σ₁ / σₗ, l = min(m, n)
05 / 11
LA 101
M07 · L03
Eckart–Young Theorem
Best Low-Rank Approximation
Truncated SVD
A_k = \sum_{i=1}^{k}\sigma_i u_i v_i^T,\quad \|A-A_k\|_2 = \sigma_{k+1}
Optimal compression
No rank-k matrix is closer to A in any unitarily invariant norm. For images, k = 50 of 1000 singular values often retains 95% of visual content.
06 / 11
LA 101
M07 · L03
Application: Pseudoinverse
Solve Any Linear System
Moore–Penrose Pseudoinverse
A^+ = V\Sigma^+ U^T,\quad x^* = A^+ b
Minimum-norm least-squares
x* = A⁺b minimizes ‖Ax − b‖² and, among all minimizers, picks the one with smallest ‖x‖. Works for any A — square, tall, wide, singular.
07 / 11
LA 101
M07 · L03
Applications
SVD Powers Data Science
- PCA: singular vectors = principal components; σᵢ² ∝ variance explained
- Recommender systems: low-rank SVD fills missing ratings
- Latent semantic analysis: topic modeling from term–document matrices
- Image compression: store only k singular values/vectors
- Noise reduction: zero out small singular values
- Numerical rank: count σᵢ > ε · σ₁
08 / 11
LA 101
M07 · L03
Computing SVD
Bidiagonalize, Then Iterate
Phase 1: reduce A to bidiagonal form B via Householder reflections — O(mn²) flops. Phase 2: apply the Golub–Reinsch QR-like iteration to find singular values of B — O(n²) per step.
Randomized SVD
For large matrices, project to a random subspace first. Finds rank-k SVD in O(mn log k) — ~100× faster than full SVD for k ≪ min(m, n).
10 / 11
LA 101
M07 · L03
Module 7 Continues
SVD: The Universal Factorization
A = UΣVᵀ works for any matrix. It exposes rank, condition number, all four subspaces, and the optimal low-rank approximation. The pseudoinverse A⁺ = VΣ⁺Uᵀ solves any system in the minimum-norm least-squares sense.
Module 7: Matrix Decompositions
SVD — Done ✓
11 / 11