LA 101
M08 · L03
Module 8: Vector Calculus Connections
Linear Algebra in Machine Learning
Every ML model is a linear algebra computation in disguise — feature matrices encode data, normal equations solve regression in closed form, and weight matrices chain together in neural networks to learn rich representations.
01 / 11
LA 101
M08 · L03
Feature Matrices
Data as a Matrix
n × p
Design matrix shape
rank(X)
Independent features
κ(XᵀX)
Conditioning
PCA connection
SVD of X = UΣVᵀ gives principal components (columns of V). The top r right singular vectors capture maximum variance. XV_r projects to r-dimensional space — dimensionality reduction as linear projection.
02 / 11
LA 101
M08 · L03
Linear Regression
Normal Equations
OLS Closed Form
\mathbf{w}^* = (X^T X)^{-1} X^T \mathbf{y}
Geometric meaning
w* is the projection of y onto col(X). Residuals y − Xw* are orthogonal to every column of X. In practice: use QR of X for numerical stability (avoids squaring the condition number).
03 / 11
LA 101
M08 · L03
Regularization
Ridge Regression
- Objective: min ‖Xw − y‖² + λ‖w‖²
- Solution: w* = (XᵀX + λI)⁻¹Xᵀy
- Effect: adds λ to every eigenvalue of XᵀX
- Condition number: (λ_max + λ)/(λ_min + λ) → improves as λ↑
- Trade-off: higher λ = more bias, less variance
04 / 11
LA 101
M08 · L03
Logistic Regression
Convex Classification
Logistic Loss
L(\mathbf{w}) = -\sum_i \left[y_i \log \sigma_i + (1-y_i)\log(1-\sigma_i)\right]
Hessian H = XᵀDX (D diagonal positive) is PSD → convex loss. Newton step = IRLS: solve (XᵀDX)δ = −∇L. Converges quadratically in ~5–10 iterations. SGD for large n.
05 / 11
LA 101
M08 · L03
Neural Networks
Chains of Matrix Multiply
Layer Forward Pass
\mathbf{z}^{(\ell)} = W_\ell \mathbf{a}^{(\ell-1)} + \mathbf{b}_\ell,\quad \mathbf{a}^{(\ell)}=\sigma(\mathbf{z}^{(\ell)})
In batch mode
Input A⁽⁰⁾ ∈ ℝⁿˣᵈ. Each layer: Z = A W ᵀ + b (matrix multiply). On GPU: GEMM kernel dominates compute. Entire LLM forward pass = sequence of matrix multiplications.
06 / 11
LA 101
M08 · L03
Backpropagation
Chain Rule in Matrix Form
- Weight grad: ∂L/∂W_ℓ = (δ^(ℓ))ᵀ A^(ℓ-1)
- Bias grad: ∂L/∂b_ℓ = sum of δ^(ℓ) over batch
- Downstream: δ^(ℓ-1) = δ^(ℓ) W_ℓ ⊙ σ'(Z^(ℓ-1))
- All operations: matrix products + elementwise multiply
- Autograd: PyTorch/JAX differentiate automatically
07 / 11
LA 101
M08 · L03
Attention Mechanism
Transformers as Matrix Algebra
Scaled Dot-Product Attention
\text{Attention}(Q,K,V)=\text{softmax}\!\left(\tfrac{QK^T}{\sqrt{d_k}}\right)V
What this does
QKᵀ ∈ ℝⁿˣⁿ is the score matrix — how much each token attends to every other. Softmax normalizes rows. Output AV is a weighted sum of value vectors. Every operation: a matrix multiply.
08 / 11
LA 101
M08 · L03
Initialization & Conditioning
Weight Matrix Geometry
- Singular values of W: σᵢ control gradient amplification
- σ_max ≫ 1: exploding gradients across layers
- σ_max ≪ 1: vanishing gradients across layers
- Xavier init: preserves variance for linear activations
- He init: scales for ReLU (half units zeroed)
- Orthogonal init: all σᵢ = 1 — perfect norm preservation
10 / 11
LA 101
M08 · L03
Module 8 · Lesson 3 Complete
ML is Linear Algebra
Design matrix X encodes data. Normal equations w* = (XᵀX)⁻¹Xᵀy solve regression. Ridge adds λI to improve conditioning. Logistic Hessian XᵀDX is PSD → convex. Neural nets = matrix chains. Attention = softmax(QKᵀ/√d)V.
Module 8: Vector Calculus Connections
Gradients · Optimization · ML Applications
11 / 11