Dimensionality Reduction
1,000 features sound rich. But in high dimensions, data is almost always sparse — distances lose meaning and models struggle. Learn to compress the noise and keep the signal.
Curse of Dimensionality
As dimensions grow, volume explodes. Points scatter so far apart that nearest-neighbor search and clustering both break down. The ratio of nearest to farthest distance approaches 1 — every point looks equally near.
PCA: Maximum Variance
PCA finds the directions of greatest variance in the data — the axes along which points spread out most. Project onto the top k of these axes and you keep the most information in k dimensions.
- Each direction is a principal component
- Components are orthogonal — no redundancy
- Computed via eigenvectors of the covariance matrix
Project onto Top Eigenvectors
Center the data, compute the covariance matrix, take its top k eigenvectors. Project the data onto those vectors. Eigenvalues tell you how much variance each component captures.
Explained Variance Ratio
Plot the cumulative explained variance as you add components. Stop at the elbow — where additional components stop contributing meaningfully. Often 95% variance needs far fewer components than features.
Linear Only, No Curves
- PCA assumes a linear subspace — no manifolds, no curves
- Cannot separate clusters that lie on a curved surface
- Sensitive to scale: always standardize features first
- Solution: use nonlinear methods for complex structure
t-SNE: Neighborhood Embedding
t-SNE converts high-dimensional distances into probabilities, then finds a 2D layout that matches those probabilities. Similar points stay close; dissimilar ones are pushed apart. Reveals cluster structure invisible to PCA.
Minimize KL Divergence
t-SNE minimizes the KL divergence between high-dimensional distribution P and low-dimensional distribution Q. A Student-t kernel in Q (heavy tails) prevents crowding in the embedding.
Visualization Only, Never Preprocessing
- Axes have no absolute meaning — only cluster shape matters
- Inter-cluster distances are not preserved
- Cannot embed new points after training
- Slow: O(n² log n) — cap at ~100,000 points
- Use PCA to 50 dims first, then t-SNE
UMAP: Faster & More Faithful
- Based on Riemannian geometry & algebraic topology
- Preserves both local and global structure
- Near-linear time — scales to millions of points
- Can embed new points after training
- Key params: n_neighbors (local vs. global) and min_dist
PCA, t-SNE, or UMAP?
- PCA — preprocessing, linear structure, always first
- t-SNE — small/medium data, tight cluster visualization
- UMAP — large data, global structure, embed new points
- Workflow: PCA → 50 dims → UMAP/t-SNE → 2D plot
What You Learned
PCA is linear and always your first step. t-SNE reveals local cluster structure for visualization. UMAP is faster, preserves global structure, and scales further. Module 5 complete.