Home / LA 101 / Module 4 / Lesson 3
Stories Mode

Diagonalization

Diagonalization rewrites a matrix as A = PDP⁻¹, where D is diagonal and P collects the eigenvectors. This decomposition turns powers, exponentials, and differential equations into trivial computations — and reveals the deepest structure of a linear transformation.

~18 min read M4 · L3 Intermediate

What Is Diagonalization?

A square matrix A is diagonalizable if it can be written as A = PDP⁻¹, where D is a diagonal matrix and P is an invertible matrix. The columns of P are eigenvectors of A, and the diagonal entries of D are the corresponding eigenvalues. This factorization is called the eigendecomposition of A.

Why does this matter? Because diagonal matrices are trivially easy to work with. If A = PDP⁻¹, then A² = PD²P⁻¹, A³ = PD³P⁻¹, and in general Aᵏ = PDᵏP⁻¹. Computing Dᵏ just means raising each diagonal entry to the k-th power. No matrix multiplication chains — just scalar operations.

Eigendecomposition
A = P D P^{-1}, \quad P = [\mathbf{v}_1 \mid \mathbf{v}_2 \mid \cdots \mid \mathbf{v}_n], \quad D = \begin{bmatrix}\lambda_1 & & \\ & \ddots & \\ & & \lambda_n\end{bmatrix}
P is the matrix whose columns are eigenvectors: P = [v₁ | v₂ | … | vₙ]. D is the diagonal matrix of eigenvalues: D = diag(λ₁, λ₂, …, λₙ). P must be invertible, meaning the eigenvectors must be linearly independent.

When Is a Matrix Diagonalizable?

An n×n matrix A is diagonalizable if and only if it has n linearly independent eigenvectors. Equivalently, for every eigenvalue, its geometric multiplicity must equal its algebraic multiplicity.

The simplest sufficient condition: if A has n distinct eigenvalues, it is automatically diagonalizable — because eigenvectors from distinct eigenvalues are always linearly independent. But distinct eigenvalues are not necessary: some matrices with repeated eigenvalues are also diagonalizable (when the repeated eigenvalue has a full-dimensional eigenspace).

Non-Diagonalizable (Defective) Matrices

The matrix [[2, 1], [0, 2]] is not diagonalizable: it has a repeated eigenvalue λ = 2 with algebraic multiplicity 2 but only one linearly independent eigenvector (geometric multiplicity 1). You cannot form an invertible P from a single eigenvector. Such matrices require the Jordan normal form — a near-diagonal form with 1s on the superdiagonal in some blocks.

The Diagonalization Procedure

To diagonalize an n×n matrix A:

  1. Find all eigenvalues. Solve det(A − λI) = 0. If A has n distinct eigenvalues, proceed directly. If there are repeated eigenvalues, check that each has sufficient independent eigenvectors.
  2. Find n linearly independent eigenvectors. For each eigenvalue, find a basis for its eigenspace. Collect all basis vectors — if the total count is n, you have your set.
  3. Form P. Place the eigenvectors as columns of P, in any order (but keep track of which eigenvalue corresponds to which column).
  4. Form D. Place the corresponding eigenvalues on the diagonal of D, in the same order as the columns of P.
  5. Verify (optional). Check that AP = PD, or equivalently that A = PDP⁻¹.

A Complete Example

Let A = [[4, 1], [2, 3]]. The characteristic polynomial is det(A − λI) = (4−λ)(3−λ) − 2 = λ² − 7λ + 10 = (λ−5)(λ−2). So the eigenvalues are λ₁ = 5 and λ₂ = 2.

Eigenvector for λ₁ = 5: Solve (A − 5I)v = 0:

A − 5I = [[−1, 1], [2, −2]] → row reduce → [[1, −1], [0, 0]]. So x₁ = x₂. Eigenvector: v₁ = [1, 1]ᵀ.

Eigenvector for λ₂ = 2: Solve (A − 2I)v = 0:

A − 2I = [[2, 1], [2, 1]] → row reduce → [[2, 1], [0, 0]]. So x₁ = −x₂/2. Setting x₂ = 2: v₂ = [−1, 2]ᵀ.

The Factorization A = PDP⁻¹
A = \begin{bmatrix}4&1\\2&3\end{bmatrix} = \begin{bmatrix}1&-1\\1&2\end{bmatrix}\begin{bmatrix}5&0\\0&2\end{bmatrix}\begin{bmatrix}1&-1\\1&2\end{bmatrix}^{-1}
P = [[1, −1], [1, 2]] collects the eigenvectors as columns. D = diag(5, 2) places the eigenvalues on the diagonal. You can verify: AP = PD.

Matrix Powers via Diagonalization

The central payoff of diagonalization is computing matrix powers efficiently. Since A = PDP⁻¹:

Aᵏ = (PDP⁻¹)ᵏ = PDP⁻¹ · PDP⁻¹ · … · PDP⁻¹ = PDᵏP⁻¹

The P⁻¹ and P in the middle cancel in pairs, leaving only the outer P and P⁻¹. And Dᵏ is trivial: diag(λ₁, λ₂, …, λₙ)ᵏ = diag(λ₁ᵏ, λ₂ᵏ, …, λₙᵏ).

Matrix Power Formula
A^k = P D^k P^{-1} = P\,\text{diag}(\lambda_1^k,\,\lambda_2^k,\,\ldots,\,\lambda_n^k)\,P^{-1}
Computing A¹⁰⁰ naively requires 99 matrix multiplications. With diagonalization: form P, D, P⁻¹ once, then raise each diagonal entry to the 100th power — just n scalar operations.

Example: A¹⁰

Using the matrix A = [[4, 1], [2, 3]] from above, A¹⁰ = PD¹⁰P⁻¹. Since D = diag(5, 2), D¹⁰ = diag(5¹⁰, 2¹⁰) = diag(9765625, 1024). The formula gives exact results without iterating 10 times.

Geometric Interpretation

Diagonalization reveals what A "really does" in the eigenvector basis. Write any vector x in terms of the eigenvectors: x = c₁v₁ + c₂v₂ + … + cₙvₙ. Then:

Ax = c₁(Av₁) + c₂(Av₂) + … + cₙ(Avₙ) = c₁λ₁v₁ + c₂λ₂v₂ + … + cₙλₙvₙ

In the eigenvector basis, A simply scales each component by the corresponding eigenvalue. The transformation is "diagonal" — no mixing between directions. This change-of-basis perspective is exactly what P and P⁻¹ accomplish: P⁻¹ converts from the standard basis to the eigenvector basis, D scales, and P converts back.

Change of Basis Perspective

The formula A = PDP⁻¹ is a change-of-basis statement. P⁻¹ expresses a vector in the eigenvector coordinate system. D then applies the transformation (just scaling). P converts back to standard coordinates. Every diagonalizable matrix is "the same as" a scaling in some basis — the eigenvector basis.

Symmetric Matrices: Always Diagonalizable

Real symmetric matrices (Aᵀ = A) are always diagonalizable — a powerful guarantee. More than that, they can be diagonalized with an orthogonal matrix: A = QDQᵀ, where Q has orthonormal columns (Qᵀ = Q⁻¹). This is the Spectral Theorem.

The Spectral Theorem means symmetric matrices have three special properties simultaneously: all eigenvalues are real, all eigenvectors can be chosen to be orthogonal, and the matrix is always diagonalizable. This is why symmetric matrices appear everywhere in applications — covariance matrices, Laplacians, Hessians — they behave beautifully.

Spectral Theorem
A = Q D Q^T, \quad Q^TQ = I, \quad A^T = A
For any real symmetric matrix A, there exists an orthogonal matrix Q (QᵀQ = I) such that A = QDQᵀ. The columns of Q are orthonormal eigenvectors; D contains the real eigenvalues. This orthogonal diagonalization is unique (up to ordering and signs of columns).

Diagonalization and Differential Equations

Diagonalization is the key to solving systems of linear differential equations. Consider x'(t) = Ax where x is a vector of unknown functions. If A = PDP⁻¹, substitute u = P⁻¹x to get u'(t) = Du(t) — a decoupled system of scalar equations: uᵢ'(t) = λᵢuᵢ(t), with solution uᵢ(t) = uᵢ(0)e^{λᵢt}.

This means the solution to x'(t) = Ax is x(t) = e^{At}x(0), where the matrix exponential e^{At} = Pe^{Dt}P⁻¹ and e^{Dt} = diag(e^{λ₁t}, …, e^{λₙt}). Diagonalization transforms an n-dimensional system into n independent scalar ODEs.

Matrix Exponential
e^{At} = P\,e^{Dt}\,P^{-1} = P\,\text{diag}(e^{\lambda_1 t},\,\ldots,\,e^{\lambda_n t})\,P^{-1}
The matrix exponential e^{At} generalizes the scalar e^{at}. For diagonalizable A = PDP⁻¹, computing e^{At} reduces to computing scalar exponentials e^{λᵢt} — a massive simplification used throughout physics and engineering.

Key Takeaways

A matrix A is diagonalizable if it has n linearly independent eigenvectors, letting us write A = PDP⁻¹. P collects eigenvectors as columns; D places eigenvalues on the diagonal. This factorization makes matrix powers trivial: Aᵏ = PDᵏP⁻¹. Geometrically, diagonalization reveals A as pure scaling in the eigenvector basis. Symmetric matrices are always diagonalizable with an orthogonal P (Spectral Theorem). The matrix exponential e^{At} = Pe^{Dt}P⁻¹ solves linear differential equation systems.