ML 101
M05 · L03
Module 5

Dimensionality Reduction

1,000 features sound rich. But in high dimensions, data is almost always sparse — distances lose meaning and models struggle. Learn to compress the noise and keep the signal.

01 / 13
ML 101
M05 · L03
The Problem

Curse of Dimensionality

As dimensions grow, volume explodes. Points scatter so far apart that nearest-neighbor search and clustering both break down. The ratio of nearest to farthest distance approaches 1 — every point looks equally near.

Problem
Sparse data
Fix
Reduce dims
02 / 13
ML 101
M05 · L03
Linear Method

PCA: Maximum Variance

PCA finds the directions of greatest variance in the data — the axes along which points spread out most. Project onto the top k of these axes and you keep the most information in k dimensions.

  • Each direction is a principal component
  • Components are orthogonal — no redundancy
  • Computed via eigenvectors of the covariance matrix
03 / 13
ML 101
M05 · L03
Formulation

Project onto Top Eigenvectors

Center the data, compute the covariance matrix, take its top k eigenvectors. Project the data onto those vectors. Eigenvalues tell you how much variance each component captures.

PCA Projection
\mathbf{Z} = \mathbf{X}\mathbf{W}_k
04 / 13
ML 101
M05 · L03
Choosing K

Explained Variance Ratio

Plot the cumulative explained variance as you add components. Stop at the elbow — where additional components stop contributing meaningfully. Often 95% variance needs far fewer components than features.

Rule of Thumb
Keep components until cumulative variance ≥ 95%
05 / 13
ML 101
M05 · L03
PCA Limits

Linear Only, No Curves

  • PCA assumes a linear subspace — no manifolds, no curves
  • Cannot separate clusters that lie on a curved surface
  • Sensitive to scale: always standardize features first
  • Solution: use nonlinear methods for complex structure
06 / 13
ML 101
M05 · L03
Nonlinear

t-SNE: Neighborhood Embedding

t-SNE converts high-dimensional distances into probabilities, then finds a 2D layout that matches those probabilities. Similar points stay close; dissimilar ones are pushed apart. Reveals cluster structure invisible to PCA.

Key param
Perplexity
Output
2D / 3D
07 / 13
ML 101
M05 · L03
Cost Function

Minimize KL Divergence

t-SNE minimizes the KL divergence between high-dimensional distribution P and low-dimensional distribution Q. A Student-t kernel in Q (heavy tails) prevents crowding in the embedding.

t-SNE Cost
C = \mathrm{KL}(P \| Q) = \sum_{i \neq j} p_{ij} \log \dfrac{p_{ij}}{q_{ij}}
08 / 13
ML 101
M05 · L03
t-SNE Rules

Visualization Only, Never Preprocessing

  • Axes have no absolute meaning — only cluster shape matters
  • Inter-cluster distances are not preserved
  • Cannot embed new points after training
  • Slow: O(n² log n) — cap at ~100,000 points
  • Use PCA to 50 dims first, then t-SNE
09 / 13
ML 101
M05 · L03
Modern Standard

UMAP: Faster & More Faithful

  • Based on Riemannian geometry & algebraic topology
  • Preserves both local and global structure
  • Near-linear time — scales to millions of points
  • Can embed new points after training
  • Key params: n_neighbors (local vs. global) and min_dist
10 / 13
ML 101
M05 · L03
Which to Use

PCA, t-SNE, or UMAP?

  • PCA — preprocessing, linear structure, always first
  • t-SNE — small/medium data, tight cluster visualization
  • UMAP — large data, global structure, embed new points
  • Workflow: PCA → 50 dims → UMAP/t-SNE → 2D plot
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

PCA is linear and always your first step. t-SNE reveals local cluster structure for visualization. UMAP is faster, preserves global structure, and scales further. Module 5 complete.

Module 5 Complete
Unsupervised Learning ✓
13 / 13