What Is Diagonalization?
A square matrix A is diagonalizable if it can be written as A = PDP⁻¹, where D is a diagonal matrix and P is an invertible matrix. The columns of P are eigenvectors of A, and the diagonal entries of D are the corresponding eigenvalues. This factorization is called the eigendecomposition of A.
Why does this matter? Because diagonal matrices are trivially easy to work with. If A = PDP⁻¹, then A² = PD²P⁻¹, A³ = PD³P⁻¹, and in general Aᵏ = PDᵏP⁻¹. Computing Dᵏ just means raising each diagonal entry to the k-th power. No matrix multiplication chains — just scalar operations.
When Is a Matrix Diagonalizable?
An n×n matrix A is diagonalizable if and only if it has n linearly independent eigenvectors. Equivalently, for every eigenvalue, its geometric multiplicity must equal its algebraic multiplicity.
The simplest sufficient condition: if A has n distinct eigenvalues, it is automatically diagonalizable — because eigenvectors from distinct eigenvalues are always linearly independent. But distinct eigenvalues are not necessary: some matrices with repeated eigenvalues are also diagonalizable (when the repeated eigenvalue has a full-dimensional eigenspace).
The matrix [[2, 1], [0, 2]] is not diagonalizable: it has a repeated eigenvalue λ = 2 with algebraic multiplicity 2 but only one linearly independent eigenvector (geometric multiplicity 1). You cannot form an invertible P from a single eigenvector. Such matrices require the Jordan normal form — a near-diagonal form with 1s on the superdiagonal in some blocks.
The Diagonalization Procedure
To diagonalize an n×n matrix A:
- Find all eigenvalues. Solve det(A − λI) = 0. If A has n distinct eigenvalues, proceed directly. If there are repeated eigenvalues, check that each has sufficient independent eigenvectors.
- Find n linearly independent eigenvectors. For each eigenvalue, find a basis for its eigenspace. Collect all basis vectors — if the total count is n, you have your set.
- Form P. Place the eigenvectors as columns of P, in any order (but keep track of which eigenvalue corresponds to which column).
- Form D. Place the corresponding eigenvalues on the diagonal of D, in the same order as the columns of P.
- Verify (optional). Check that AP = PD, or equivalently that A = PDP⁻¹.
A Complete Example
Let A = [[4, 1], [2, 3]]. The characteristic polynomial is det(A − λI) = (4−λ)(3−λ) − 2 = λ² − 7λ + 10 = (λ−5)(λ−2). So the eigenvalues are λ₁ = 5 and λ₂ = 2.
Eigenvector for λ₁ = 5: Solve (A − 5I)v = 0:
A − 5I = [[−1, 1], [2, −2]] → row reduce → [[1, −1], [0, 0]]. So x₁ = x₂. Eigenvector: v₁ = [1, 1]ᵀ.
Eigenvector for λ₂ = 2: Solve (A − 2I)v = 0:
A − 2I = [[2, 1], [2, 1]] → row reduce → [[2, 1], [0, 0]]. So x₁ = −x₂/2. Setting x₂ = 2: v₂ = [−1, 2]ᵀ.
Matrix Powers via Diagonalization
The central payoff of diagonalization is computing matrix powers efficiently. Since A = PDP⁻¹:
Aᵏ = (PDP⁻¹)ᵏ = PDP⁻¹ · PDP⁻¹ · … · PDP⁻¹ = PDᵏP⁻¹
The P⁻¹ and P in the middle cancel in pairs, leaving only the outer P and P⁻¹. And Dᵏ is trivial: diag(λ₁, λ₂, …, λₙ)ᵏ = diag(λ₁ᵏ, λ₂ᵏ, …, λₙᵏ).
Example: A¹⁰
Using the matrix A = [[4, 1], [2, 3]] from above, A¹⁰ = PD¹⁰P⁻¹. Since D = diag(5, 2), D¹⁰ = diag(5¹⁰, 2¹⁰) = diag(9765625, 1024). The formula gives exact results without iterating 10 times.
Geometric Interpretation
Diagonalization reveals what A "really does" in the eigenvector basis. Write any vector x in terms of the eigenvectors: x = c₁v₁ + c₂v₂ + … + cₙvₙ. Then:
Ax = c₁(Av₁) + c₂(Av₂) + … + cₙ(Avₙ) = c₁λ₁v₁ + c₂λ₂v₂ + … + cₙλₙvₙ
In the eigenvector basis, A simply scales each component by the corresponding eigenvalue. The transformation is "diagonal" — no mixing between directions. This change-of-basis perspective is exactly what P and P⁻¹ accomplish: P⁻¹ converts from the standard basis to the eigenvector basis, D scales, and P converts back.
The formula A = PDP⁻¹ is a change-of-basis statement. P⁻¹ expresses a vector in the eigenvector coordinate system. D then applies the transformation (just scaling). P converts back to standard coordinates. Every diagonalizable matrix is "the same as" a scaling in some basis — the eigenvector basis.
Symmetric Matrices: Always Diagonalizable
Real symmetric matrices (Aᵀ = A) are always diagonalizable — a powerful guarantee. More than that, they can be diagonalized with an orthogonal matrix: A = QDQᵀ, where Q has orthonormal columns (Qᵀ = Q⁻¹). This is the Spectral Theorem.
The Spectral Theorem means symmetric matrices have three special properties simultaneously: all eigenvalues are real, all eigenvectors can be chosen to be orthogonal, and the matrix is always diagonalizable. This is why symmetric matrices appear everywhere in applications — covariance matrices, Laplacians, Hessians — they behave beautifully.
Diagonalization and Differential Equations
Diagonalization is the key to solving systems of linear differential equations. Consider x'(t) = Ax where x is a vector of unknown functions. If A = PDP⁻¹, substitute u = P⁻¹x to get u'(t) = Du(t) — a decoupled system of scalar equations: uᵢ'(t) = λᵢuᵢ(t), with solution uᵢ(t) = uᵢ(0)e^{λᵢt}.
This means the solution to x'(t) = Ax is x(t) = e^{At}x(0), where the matrix exponential e^{At} = Pe^{Dt}P⁻¹ and e^{Dt} = diag(e^{λ₁t}, …, e^{λₙt}). Diagonalization transforms an n-dimensional system into n independent scalar ODEs.
A matrix A is diagonalizable if it has n linearly independent eigenvectors, letting us write A = PDP⁻¹. P collects eigenvectors as columns; D places eigenvalues on the diagonal. This factorization makes matrix powers trivial: Aᵏ = PDᵏP⁻¹. Geometrically, diagonalization reveals A as pure scaling in the eigenvector basis. Symmetric matrices are always diagonalizable with an orthogonal P (Spectral Theorem). The matrix exponential e^{At} = Pe^{Dt}P⁻¹ solves linear differential equation systems.