LA 101
M08 · L03
Module 8: Vector Calculus Connections

Linear Algebra in Machine Learning

Every ML model is a linear algebra computation in disguise — feature matrices encode data, normal equations solve regression in closed form, and weight matrices chain together in neural networks to learn rich representations.

01 / 11
LA 101
M08 · L03
Feature Matrices

Data as a Matrix

n × p
Design matrix shape
rank(X)
Independent features
κ(XᵀX)
Conditioning
PCA connection
SVD of X = UΣVᵀ gives principal components (columns of V). The top r right singular vectors capture maximum variance. XV_r projects to r-dimensional space — dimensionality reduction as linear projection.
02 / 11
LA 101
M08 · L03
Linear Regression

Normal Equations

OLS Closed Form
\mathbf{w}^* = (X^T X)^{-1} X^T \mathbf{y}
Geometric meaning
w* is the projection of y onto col(X). Residuals y − Xw* are orthogonal to every column of X. In practice: use QR of X for numerical stability (avoids squaring the condition number).
03 / 11
LA 101
M08 · L03
Regularization

Ridge Regression

  • Objective: min ‖Xw − y‖² + λ‖w‖²
  • Solution: w* = (XᵀX + λI)⁻¹Xᵀy
  • Effect: adds λ to every eigenvalue of XᵀX
  • Condition number: (λ_max + λ)/(λ_min + λ) → improves as λ↑
  • Trade-off: higher λ = more bias, less variance
04 / 11
LA 101
M08 · L03
Logistic Regression

Convex Classification

Logistic Loss
L(\mathbf{w}) = -\sum_i \left[y_i \log \sigma_i + (1-y_i)\log(1-\sigma_i)\right]

Hessian H = XᵀDX (D diagonal positive) is PSD → convex loss. Newton step = IRLS: solve (XᵀDX)δ = −∇L. Converges quadratically in ~5–10 iterations. SGD for large n.

05 / 11
LA 101
M08 · L03
Neural Networks

Chains of Matrix Multiply

Layer Forward Pass
\mathbf{z}^{(\ell)} = W_\ell \mathbf{a}^{(\ell-1)} + \mathbf{b}_\ell,\quad \mathbf{a}^{(\ell)}=\sigma(\mathbf{z}^{(\ell)})
In batch mode
Input A⁽⁰⁾ ∈ ℝⁿˣᵈ. Each layer: Z = A W ᵀ + b (matrix multiply). On GPU: GEMM kernel dominates compute. Entire LLM forward pass = sequence of matrix multiplications.
06 / 11
LA 101
M08 · L03
Backpropagation

Chain Rule in Matrix Form

  • Weight grad: ∂L/∂W_ℓ = (δ^(ℓ))ᵀ A^(ℓ-1)
  • Bias grad: ∂L/∂b_ℓ = sum of δ^(ℓ) over batch
  • Downstream: δ^(ℓ-1) = δ^(ℓ) W_ℓ ⊙ σ'(Z^(ℓ-1))
  • All operations: matrix products + elementwise multiply
  • Autograd: PyTorch/JAX differentiate automatically
07 / 11
LA 101
M08 · L03
Attention Mechanism

Transformers as Matrix Algebra

Scaled Dot-Product Attention
\text{Attention}(Q,K,V)=\text{softmax}\!\left(\tfrac{QK^T}{\sqrt{d_k}}\right)V
What this does
QKᵀ ∈ ℝⁿˣⁿ is the score matrix — how much each token attends to every other. Softmax normalizes rows. Output AV is a weighted sum of value vectors. Every operation: a matrix multiply.
08 / 11
LA 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

09 / 11
LA 101
M08 · L03
Initialization & Conditioning

Weight Matrix Geometry

  • Singular values of W: σᵢ control gradient amplification
  • σ_max ≫ 1: exploding gradients across layers
  • σ_max ≪ 1: vanishing gradients across layers
  • Xavier init: preserves variance for linear activations
  • He init: scales for ReLU (half units zeroed)
  • Orthogonal init: all σᵢ = 1 — perfect norm preservation
10 / 11
LA 101
M08 · L03
Module 8 · Lesson 3 Complete

ML is Linear Algebra

Design matrix X encodes data. Normal equations w* = (XᵀX)⁻¹Xᵀy solve regression. Ridge adds λI to improve conditioning. Logistic Hessian XᵀDX is PSD → convex. Neural nets = matrix chains. Attention = softmax(QKᵀ/√d)V.

Module 8: Vector Calculus Connections
Gradients · Optimization · ML Applications
11 / 11