Home / LA 101 / Module 5 / Lesson 1
Stories Mode

Orthogonal Vectors and Projections

When two vectors meet at a right angle, something powerful happens — they share no information. Orthogonality is the geometric foundation of projections, least squares, and the Fourier transform.

~15 min read M5 · L1 Intermediate

When Vectors Are Perpendicular

In everyday geometry, two lines are perpendicular when they meet at a 90° angle. In linear algebra, the same idea generalizes to any number of dimensions using the dot product. Two vectors v and w are orthogonal if their dot product is zero:

Orthogonality Condition
\mathbf{v} \cdot \mathbf{w} = \mathbf{v}^T\mathbf{w} = 0
Two vectors are orthogonal if and only if their dot product is zero. Geometrically, this means they meet at a right angle. Note that the zero vector is orthogonal to every vector.

Why does the dot product measure angles? Because the dot product of v and w equals ||v|| ||w|| cos θ, where θ is the angle between them. When θ = 90°, cos θ = 0, so the dot product vanishes. Orthogonality is the algebraic certificate of perpendicularity.

A Concrete Example

The standard basis vectors in ℝ² are orthogonal: e₁ = [1, 0]ᵀ and e₂ = [0, 1]ᵀ. Their dot product is 1·0 + 0·1 = 0. More interestingly, [3, −4]ᵀ and [4, 3]ᵀ are orthogonal: 3·4 + (−4)·3 = 12 − 12 = 0. You can check this without drawing a picture — the algebra reveals the geometry.

Pythagoras in Higher Dimensions

If v and w are orthogonal, then ||v + w||² = ||v||² + ||w||². This is the Pythagorean theorem for vectors — it holds in any dimension, not just 2D or 3D. Orthogonality generalizes right triangles to n-dimensional space.

Orthogonal Complements

The orthogonal complement of a subspace W, written W⊥ (pronounced "W perp"), is the set of all vectors that are orthogonal to every vector in W:

Orthogonal Complement
W^\perp = \{\mathbf{x} \in \mathbb{R}^n : \mathbf{x} \cdot \mathbf{w} = 0 \text{ for all } \mathbf{w} \in W\}
W⊥ is itself a subspace. Together W and W⊥ cover all of ℝⁿ without overlap: dim(W) + dim(W⊥) = n.

Think of a plane through the origin in ℝ³. Its orthogonal complement is the line through the origin perpendicular to that plane — the normal line. Every vector in ℝ³ can be split uniquely into a piece lying in the plane and a piece along the normal. This splitting is the heart of projection.

The Four Fundamental Subspaces

The orthogonal complement connects directly to the structure of any matrix A. The null space of A and the row space of A are orthogonal complements in ℝⁿ. The left null space and the column space are orthogonal complements in ℝᵐ. These four subspaces form the "big picture" of linear algebra, and orthogonality is the glue that holds them together.

Projection onto a Line

Given a vector b and a direction vector a, the projection of b onto the line through a is the closest point on that line to b. It is the shadow of b cast perpendicularly onto the line.

To find it, we want a scalar c such that ca is closest to b. The error vector b − ca must be perpendicular to a:

aᵀ(b − ca) = 0 → c = aᵀb / aᵀa

So the projection vector is:

Projection onto a Line
\text{proj}_{\mathbf{a}}\mathbf{b} = \frac{\mathbf{a}^T\mathbf{b}}{\mathbf{a}^T\mathbf{a}}\,\mathbf{a}
The scalar factor (aᵀb / aᵀa) is how far along a the projection falls. The projection matrix P = aaᵀ / aᵀa projects any vector onto the line spanned by a.

The matrix P = aaᵀ / aᵀa is the projection matrix onto the line through a. It has two key properties: P² = P (projecting twice gives the same result as projecting once) and Pᵀ = P (it is symmetric). Any matrix with these two properties is a projection matrix.

Numerical Example

Project b = [1, 2]ᵀ onto a = [3, 4]ᵀ. The scalar is c = (3·1 + 4·2) / (3² + 4²) = (3 + 8) / 25 = 11/25. So the projection is (11/25) · [3, 4]ᵀ = [33/25, 44/25]ᵀ ≈ [1.32, 1.76]ᵀ. The error vector is [1, 2]ᵀ − [1.32, 1.76]ᵀ = [−0.32, 0.24]ᵀ. Check: [3, 4]ᵀ · [−0.32, 0.24]ᵀ = −0.96 + 0.96 = 0. The error is indeed orthogonal to a.

Projection onto a Subspace

The idea extends naturally. Suppose we have a subspace spanned by columns of a matrix A (with linearly independent columns). The projection of b onto the column space of A is the vector in col(A) closest to b. The error must be perpendicular to every column of A, which means it lies in the left null space of A:

Projection onto a Subspace
P = A(A^TA)^{-1}A^T, \quad \hat{\mathbf{x}} = (A^TA)^{-1}A^T\mathbf{b}
The projection matrix P = A(AᵀA)⁻¹Aᵀ projects any vector onto the column space of A. When A has orthonormal columns (QᵀQ = I), this simplifies beautifully to P = QQᵀ.

The formula AᵀAx̂ = Aᵀb is called the normal equation. Its solution x̂ gives the coordinates of the projection in terms of the columns of A. This is the gateway to least squares — the topic of Lesson 5.4.

Orthogonal Decomposition

The most elegant consequence of projection is the orthogonal decomposition theorem: every vector b in ℝⁿ can be written uniquely as the sum of a component in a subspace W and a component in its orthogonal complement W⊥:

Orthogonal Decomposition
\mathbf{b} = \underbrace{P\mathbf{b}}_{\mathbf{p} \,\in\, W} + \underbrace{(I-P)\mathbf{b}}_{\mathbf{e} \,\in\, W^\perp}
Every vector splits uniquely into a projection p (in W) and an error e (in W⊥). The two pieces are orthogonal: p · e = 0. This is the fundamental theorem of orthogonal projection.

The component p = Pb is the projection onto W. The component e = b − p = (I − P)b is the residual, sometimes called the error or rejection. Notice that I − P is itself a projection matrix — it projects onto W⊥. The two projections P and I − P partition the identity: P + (I − P) = I.

Why This Matters

The orthogonal decomposition is not just a geometric curiosity. It is the engine behind: least squares fitting (find the closest point in a subspace), Gram-Schmidt orthogonalization (build an orthogonal basis step by step), QR decomposition (factor any matrix into orthogonal and triangular parts), and the Fourier transform (decompose a signal into orthogonal frequency components). Orthogonality is the organizing principle of computational linear algebra.

Connection to Signal Processing

In signal processing, projecting a signal onto a subspace spanned by sinusoids gives you the Fourier coefficients. Each coefficient measures how much of that frequency is present in the signal. The residual is the part of the signal that the sinusoids cannot explain — it is orthogonal to every frequency component you included.


Key Takeaways

Two vectors are orthogonal when their dot product is zero — they share no directional information. The orthogonal complement W⊥ contains all vectors perpendicular to a subspace W, and dim(W) + dim(W⊥) = n. The projection of b onto a line through a is proj = (aᵀb / aᵀa)a. For a subspace spanned by A, the projection matrix is P = A(AᵀA)⁻¹Aᵀ. Every vector decomposes uniquely as b = p + e where p is in W and e is in W⊥ — and p · e = 0.