Matrix Factorization
Large data matrices hide structure no single entry reveals. Matrix factorization — NMF, SVD-based collaborative filtering, LSA — decomposes a matrix into interpretable parts, uncovering latent factors that explain the observed data.
Parts-Based Decomposition
Given V ≥ 0, find W ≥ 0, H ≥ 0 with V ≈ WH. The non-negativity forces additive, parts-based factors — each column of W is a "part", each row of H is a mixture weight. No cancellation allowed.
Multiplicative Updates
NMF alternates between updating H and W using ratio-based gradient steps that guarantee non-negativity throughout. Converges to a local minimum — the problem is non-convex in (W,H) jointly.
Parts vs. Holistic
PCA allows negative coefficients — components can cancel, producing holistic representations. NMF only adds, never subtracts. For images this means facial features (eyes, nose, mouth), not eigen-face blends. Far more interpretable in human terms.
Collaborative Filtering
Ratings matrix R is low-rank: a few latent factors (genres, styles) explain most variance. Learn user vectors pᵢ and item vectors qⱼ in a shared latent space. Predicted rating: r̂ᵢⱼ = pᵢᵀqⱼ.
Users, Items, and Similarity
After training, pᵢ and qⱼ live in the same k-dimensional space. Similar items cluster — even with no shared explicit features. Similar users cluster too. The dot product pᵢᵀqⱼ measures alignment: large value → predicted high rating.
Optimal Low-Rank Approximation
The best rank-k approximation to A (in Frobenius and spectral norm) is the truncated SVD: keep only the k largest singular values. Error = σₖ₊₁² + … + σᵣ². No other rank-k matrix gets closer.
SVD on Text
Build a TF-IDF term-document matrix A. Truncate its SVD to rank k. The resulting concept vectors capture synonymy (same-meaning words cluster) and survive polysemy. Document similarity = cosine of concept vectors — even documents sharing no words can be similar.
- Rows of UₖΣₖ: concept vectors for terms
- Columns of ΣₖVₖᵀ: concept vectors for documents
- k components = k latent "topics"
GloVe and PMI Factorization
GloVe explicitly factorizes the log co-occurrence matrix: wᵢᵀw̃ⱼ ≈ log Xᵢⱼ. Word2Vec implicitly does the same for the shifted PMI matrix. Result: semantic analogies as geometry — king − man + woman ≈ queen in embedding space.
Key Takeaways
- NMF: V ≈ WH with W,H ≥ 0 → parts-based, interpretable factors
- Collaborative filtering: low-rank R ≈ PQᵀ; predict rᵢⱼ = pᵢᵀqⱼ
- Eckart–Young: truncated SVD is the optimal rank-k approximation
- LSA: SVD on TF-IDF matrix captures synonymy and latent topics
- Word embeddings (GloVe/W2V): implicitly factorize log co-occurrence — analogy = geometry
M11-L3: Neural Network Fundamentals
We've seen how matrix factorization discovers hidden structure. Next: how neural networks chain matrix multiplications — forward propagation, backpropagation as Jacobian chain rule, and the attention mechanism as a QKV matrix operation.