LA 101
M10 · L03
Module 10: Signal Processing Applications

PCA and Dimensionality Reduction

Data has structure. PCA finds it by eigendecomposing the covariance matrix — extracting the directions of maximum variance, compressing representations, and denoising high-dimensional signals.

01 / 12
LA 101
M10 · L03
The Setup

Data as a Cloud in High Dimensions

N observations, each with d features, form a cloud of N points in ℝᵈ. Natural data is highly redundant — the cloud lives near a much lower-dimensional subspace. PCA finds that subspace.

The Question
What is the best k-dimensional linear subspace to represent the data — minimizing reconstruction error and maximizing preserved variance?
02 / 12
LA 101
M10 · L03
Step 1: Center, then Compute

The Covariance Matrix

Subtract the mean of each feature to get X̃. The sample covariance matrix encodes all pairwise feature correlations:

Sample Covariance
C=\tfrac{1}{N-1}\tilde{X}^T\tilde{X}

C is symmetric and positive semidefinite. High off-diagonal values mean redundancy — compressibility.

03 / 12
LA 101
M10 · L03
Step 2: Eigendecompose

Spectral Decomposition

The spectral theorem gives a complete orthonormal eigenbasis:

Eigendecomposition
C=Q\Lambda Q^T,\quad\Lambda=\mathrm{diag}(\lambda_1,\ldots,\lambda_d)

Each eigenvalue λₖ = variance along eigenvector qₖ. The eigenvectors are the principal axes of the data ellipsoid.

04 / 12
LA 101
M10 · L03
Step 3: Project

PCA Projection onto k Components

Keep the top k eigenvectors. The low-dimensional representation is:

PCA Scores
\mathbf{z}=Q_k^T\,\tilde{\mathbf{x}}\;\in\;\mathbb{R}^k

This is the optimal rank-k encoder — among all linear projections onto k dimensions, this minimizes reconstruction error (Eckart-Young theorem).

05 / 12
LA 101
M10 · L03
How Many Components?

Explained Variance Ratio

The fraction of total variance captured by the top k components:

EVR
\mathrm{EVR}(k)=\frac{\sum_{j=1}^k\lambda_j}{\sum_{j=1}^d\lambda_j}

Choose k so EVR(k) ≥ 0.95. For images, speech, and sensor data, 95% variance often fits in <5% of dimensions.

06 / 12
LA 101
M10 · L03
Visual Guide

The Scree Plot

Plot eigenvalues in decreasing order. Look for the "elbow" — the point where the curve bends from steep to flat. The elbow reveals the intrinsic dimensionality of the data.

  • Before the elbow: signal components — large, widely varying eigenvalues
  • After the elbow: noise floor — small, nearly constant eigenvalues
  • No elbow: data is genuinely high-dimensional or poorly structured
  • Multiple elbows: hierarchical structure at different scales
07 / 12
LA 101
M10 · L03
Applications

Compression and Visualization

PCA powers two immediate applications: compression (keep k ≪ d coefficients per sample) and visualization (project to 2D or 3D to see cluster structure).

k=50
of 784 pixels
95%
variance kept
16×
compression
08 / 12
LA 101
M10 · L03
Signal Processing

Noise Reduction via Subspace Projection

For noisy data x̃ = s + n (rank-k signal + white noise σ²I), the top k eigenvalues exceed σ². Project onto the signal subspace to suppress noise:

Subspace Denoiser
\hat{\mathbf{s}}=Q_k Q_k^T\,\tilde{\mathbf{x}}

This hard-threshold denoiser is optimal in Frobenius norm for known-rank signals (Eckart-Young).

09 / 12
LA 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 12
LA 101
Key Takeaways
Summary

Key Takeaways

  • PCA eigendecomposes the covariance C = QΛQᵀ; eigenvalues = variance along principal directions
  • Top k eigenvectors give the optimal rank-k linear encoder (Eckart-Young theorem)
  • EVR(k) = (λ₁+…+λₖ)/tr(C) measures how much variance is preserved
  • Scree plot elbow reveals intrinsic dimensionality; use 95% EVR as default threshold
  • Signal subspace projection QₖQₖᵀx̃ optimally denoises rank-k signals in white noise
11 / 12
LA 101
Module Complete
Module 10 Complete

Linear Algebra in Signal Processing

You have seen linear algebra meet real signals: convolution as matrix multiplication, optimal filter design via least squares and Wiener-Hopf, and now PCA turning the eigendecomposition into a data compression and denoising machine.

Course Overview
12 / 12