LA 101
M07 · L04
Module 7: Matrix Decompositions
Cholesky Decomposition
A = LLᵀ for symmetric positive definite matrices — twice as fast as LU, no pivoting needed, the foundation of statistical simulation and Kalman filtering.
01 / 11
LA 101
M07 · L04
Symmetric Positive Definite Matrices
When Is A Positive Definite?
Aᵀ = A
Symmetric
xᵀAx > 0
Positive definite
λᵢ > 0
All eigenvalues positive
Common sources
Normal equations AᵀA (full-rank A), covariance matrices in statistics, stiffness matrices in FEM, kernel matrices in Gaussian processes.
02 / 11
LA 101
M07 · L04
The Factorization
A = LLᵀ
Cholesky Decomposition
A = LL^T,\quad L \text{ lower triangular},\quad \ell_{ii} > 0
Unique
For any SPD matrix, there exists exactly one lower triangular L with positive diagonal entries satisfying A = LLᵀ. It is the triangular "square root" of A.
03 / 11
LA 101
M07 · L04
The Algorithm
Column by Column
Cholesky Recurrence
\ell_{jj}=\sqrt{a_{jj}-\textstyle\sum_{k<j}\ell_{jk}^2},\quad \ell_{ij}=\frac{a_{ij}-\sum_{k<j}\ell_{ik}\ell_{jk}}{\ell_{jj}}
If the argument to √ is ≤ 0, A is not positive definite — the algorithm doubles as an SPD test.
04 / 11
LA 101
M07 · L04
Efficiency & Stability
Half the Cost of LU
- n³/6 flops — half of LU's n³/3
- No pivoting — positive definiteness guarantees stability
- Half the storage — only lower triangle needed
- Growth factor ≤ 1 — unconditionally backward-stable
- SPD test built-in — failure = not positive definite
05 / 11
LA 101
M07 · L04
Solving Linear Systems
Two Triangular Solves
Ax = b via A = LLᵀ
Ly=b\;(\text{forward})\;\Rightarrow\;L^Tx=y\;(\text{back})
Determinant
det(A) = (ℓ₁₁ · ℓ₂₂ · ⋯ · ℓₙₙ)². Log-det for Gaussian likelihoods: log det(A) = 2 · Σᵢ log ℓᵢᵢ.
06 / 11
LA 101
M07 · L04
Application: Statistical Simulation
Sample Correlated Random Vectors
Multivariate Gaussian sampling
\Sigma=LL^T,\;z\sim\mathcal{N}(0,I)\;\Rightarrow\;x=\mu+Lz\sim\mathcal{N}(\mu,\Sigma)
Why it works
E[LzᵀzᵀLᵀ] = L·I·Lᵀ = LLᵀ = Σ. The Cholesky factor L is the "colorization" operator that introduces the desired correlations.
07 / 11
LA 101
M07 · L04
Applications
Where Cholesky Appears
- Normal equations: solve AᵀAx = Aᵀb for least squares
- Kalman filter: square-root variants propagate L of covariance P
- Gaussian process: solve Kα = y and compute log det(K)
- Newton's method: solve Hessian systems at each iteration
- FEM: sparse Cholesky for large stiffness matrices
- Bayesian inference: posterior covariance sampling
08 / 11
LA 101
M07 · L04
Comparison
Cholesky vs. Other Methods
- vs. LU: 2× faster, no pivoting, SPD only
- vs. QR: QR avoids κ² conditioning for least squares; Cholesky is faster
- vs. SVD: SVD reveals full spectrum; Cholesky just solves fast
- vs. Eigendecomp: eigendecomp needs all eigenvalues; Cholesky just needs A = LLᵀ
Rule of thumb
If A is SPD and you just need to solve Ax = b — use Cholesky. It's the fastest, most stable choice.
10 / 11
LA 101
M07 · L04
Module 7 Complete
Cholesky: The SPD Specialist
A = LLᵀ — faster, simpler, and more stable than LU for the most important class of matrices. From statistics to Kalman filters, Cholesky is the workhorse of SPD computation.
Module 7: Matrix Decompositions
LU · QR · SVD · Cholesky ✓
11 / 11