LA 101
M07 · L04
Module 7: Matrix Decompositions

Cholesky Decomposition

A = LLᵀ for symmetric positive definite matrices — twice as fast as LU, no pivoting needed, the foundation of statistical simulation and Kalman filtering.

01 / 11
LA 101
M07 · L04
Symmetric Positive Definite Matrices

When Is A Positive Definite?

Aᵀ = A
Symmetric
xᵀAx > 0
Positive definite
λᵢ > 0
All eigenvalues positive
Common sources
Normal equations AᵀA (full-rank A), covariance matrices in statistics, stiffness matrices in FEM, kernel matrices in Gaussian processes.
02 / 11
LA 101
M07 · L04
The Factorization

A = LLᵀ

Cholesky Decomposition
A = LL^T,\quad L \text{ lower triangular},\quad \ell_{ii} > 0
Unique
For any SPD matrix, there exists exactly one lower triangular L with positive diagonal entries satisfying A = LLᵀ. It is the triangular "square root" of A.
03 / 11
LA 101
M07 · L04
The Algorithm

Column by Column

Cholesky Recurrence
\ell_{jj}=\sqrt{a_{jj}-\textstyle\sum_{k<j}\ell_{jk}^2},\quad \ell_{ij}=\frac{a_{ij}-\sum_{k<j}\ell_{ik}\ell_{jk}}{\ell_{jj}}

If the argument to √ is ≤ 0, A is not positive definite — the algorithm doubles as an SPD test.

04 / 11
LA 101
M07 · L04
Efficiency & Stability

Half the Cost of LU

  • n³/6 flops — half of LU's n³/3
  • No pivoting — positive definiteness guarantees stability
  • Half the storage — only lower triangle needed
  • Growth factor ≤ 1 — unconditionally backward-stable
  • SPD test built-in — failure = not positive definite
05 / 11
LA 101
M07 · L04
Solving Linear Systems

Two Triangular Solves

Ax = b via A = LLᵀ
Ly=b\;(\text{forward})\;\Rightarrow\;L^Tx=y\;(\text{back})
Determinant
det(A) = (ℓ₁₁ · ℓ₂₂ · ⋯ · ℓₙₙ)². Log-det for Gaussian likelihoods: log det(A) = 2 · Σᵢ log ℓᵢᵢ.
06 / 11
LA 101
M07 · L04
Application: Statistical Simulation

Sample Correlated Random Vectors

Multivariate Gaussian sampling
\Sigma=LL^T,\;z\sim\mathcal{N}(0,I)\;\Rightarrow\;x=\mu+Lz\sim\mathcal{N}(\mu,\Sigma)
Why it works
E[LzᵀzᵀLᵀ] = L·I·Lᵀ = LLᵀ = Σ. The Cholesky factor L is the "colorization" operator that introduces the desired correlations.
07 / 11
LA 101
M07 · L04
Applications

Where Cholesky Appears

  • Normal equations: solve AᵀAx = Aᵀb for least squares
  • Kalman filter: square-root variants propagate L of covariance P
  • Gaussian process: solve Kα = y and compute log det(K)
  • Newton's method: solve Hessian systems at each iteration
  • FEM: sparse Cholesky for large stiffness matrices
  • Bayesian inference: posterior covariance sampling
08 / 11
LA 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

09 / 11
LA 101
M07 · L04
Comparison

Cholesky vs. Other Methods

  • vs. LU: 2× faster, no pivoting, SPD only
  • vs. QR: QR avoids κ² conditioning for least squares; Cholesky is faster
  • vs. SVD: SVD reveals full spectrum; Cholesky just solves fast
  • vs. Eigendecomp: eigendecomp needs all eigenvalues; Cholesky just needs A = LLᵀ
Rule of thumb
If A is SPD and you just need to solve Ax = b — use Cholesky. It's the fastest, most stable choice.
10 / 11
LA 101
M07 · L04
Module 7 Complete

Cholesky: The SPD Specialist

A = LLᵀ — faster, simpler, and more stable than LU for the most important class of matrices. From statistics to Kalman filters, Cholesky is the workhorse of SPD computation.

Module 7: Matrix Decompositions
LU · QR · SVD · Cholesky ✓
11 / 11