PCA and Dimensionality Reduction
Data has structure. PCA finds it by eigendecomposing the covariance matrix — extracting the directions of maximum variance, compressing representations, and denoising high-dimensional signals.
Data as a Cloud in High Dimensions
N observations, each with d features, form a cloud of N points in ℝᵈ. Natural data is highly redundant — the cloud lives near a much lower-dimensional subspace. PCA finds that subspace.
The Covariance Matrix
Subtract the mean of each feature to get X̃. The sample covariance matrix encodes all pairwise feature correlations:
C is symmetric and positive semidefinite. High off-diagonal values mean redundancy — compressibility.
Spectral Decomposition
The spectral theorem gives a complete orthonormal eigenbasis:
Each eigenvalue λₖ = variance along eigenvector qₖ. The eigenvectors are the principal axes of the data ellipsoid.
PCA Projection onto k Components
Keep the top k eigenvectors. The low-dimensional representation is:
This is the optimal rank-k encoder — among all linear projections onto k dimensions, this minimizes reconstruction error (Eckart-Young theorem).
Explained Variance Ratio
The fraction of total variance captured by the top k components:
Choose k so EVR(k) ≥ 0.95. For images, speech, and sensor data, 95% variance often fits in <5% of dimensions.
The Scree Plot
Plot eigenvalues in decreasing order. Look for the "elbow" — the point where the curve bends from steep to flat. The elbow reveals the intrinsic dimensionality of the data.
- Before the elbow: signal components — large, widely varying eigenvalues
- After the elbow: noise floor — small, nearly constant eigenvalues
- No elbow: data is genuinely high-dimensional or poorly structured
- Multiple elbows: hierarchical structure at different scales
Compression and Visualization
PCA powers two immediate applications: compression (keep k ≪ d coefficients per sample) and visualization (project to 2D or 3D to see cluster structure).
Noise Reduction via Subspace Projection
For noisy data x̃ = s + n (rank-k signal + white noise σ²I), the top k eigenvalues exceed σ². Project onto the signal subspace to suppress noise:
This hard-threshold denoiser is optimal in Frobenius norm for known-rank signals (Eckart-Young).
Key Takeaways
- PCA eigendecomposes the covariance C = QΛQᵀ; eigenvalues = variance along principal directions
- Top k eigenvectors give the optimal rank-k linear encoder (Eckart-Young theorem)
- EVR(k) = (λ₁+…+λₖ)/tr(C) measures how much variance is preserved
- Scree plot elbow reveals intrinsic dimensionality; use 95% EVR as default threshold
- Signal subspace projection QₖQₖᵀx̃ optimally denoises rank-k signals in white noise
Linear Algebra in Signal Processing
You have seen linear algebra meet real signals: convolution as matrix multiplication, optimal filter design via least squares and Wiener-Hopf, and now PCA turning the eigendecomposition into a data compression and denoising machine.