LA 101
M08 · L02
Module 8: Vector Calculus Connections
Optimization with Linear Algebra
Gradient descent navigates loss landscapes step by step. Newton's method uses curvature for faster convergence. Lagrange multipliers enforce constraints. Together they make linear algebra the engine of all modern optimization.
01 / 11
LA 101
M08 · L02
Gradient Descent
Step in the Downhill Direction
−α∇f
Update direction
α
Step size (learning rate)
‖∇f‖→0
Convergence criterion
Convergence Rate
Convergence speed is governed by the condition number κ = λ_max/λ_min of the Hessian. High κ means slow convergence — the reason preconditioning matters.
02 / 11
LA 101
M08 · L02
The Update Rule
GD Equation
Gradient Descent Update
\mathbf{x}_{k+1} = \mathbf{x}_k - \alpha\,\nabla f(\mathbf{x}_k)
Step size intuition
α too large → overshoot and diverge. α too small → slow convergence. Optimal α ≈ 2/(λ_max + λ_min) for a quadratic. Line search finds the best α per iteration.
03 / 11
LA 101
M08 · L02
Loss Landscape
Convex vs Non-Convex
- Convex: single bowl — any local min is global min
- Non-convex: rugged — multiple local minima possible
- Saddle points: ∇f = 0 but H indefinite — can trap GD
- Flat plateaus: ‖∇f‖ ≈ 0 everywhere — very slow progress
- SGD advantage: noise helps escape saddle points
04 / 11
LA 101
M08 · L02
Newton's Method
Use the Hessian
- Second-order method: uses H(x_k) at each step
- Quadratic convergence: correct digits double per step
- Cost: O(n³) per step — infeasible for large n
- Newton direction: δ = −H⁻¹∇f solves the local quadratic exactly
- Quasi-Newton (L-BFGS): approximate H⁻¹ from gradient history
- Safeguard: if H indefinite, add λI to force PD
05 / 11
LA 101
M08 · L02
Newton's Update Rule
Curvature-Aware Step
Newton Update
\mathbf{x}_{k+1} = \mathbf{x}_k - H(\mathbf{x}_k)^{-1}\nabla f(\mathbf{x}_k)
vs gradient descent
GD uses −α∇f (ignores curvature). Newton uses −H⁻¹∇f (accounts for curvature). In steep directions the step shrinks; in flat directions it grows. This is why Newton converges so fast near a minimum.
06 / 11
LA 101
M08 · L02
Convexity
PD Hessian → Global Min
Convexity Condition
f(\lambda\mathbf{x}+(1-\lambda)\mathbf{y}) \leq \lambda f(\mathbf{x})+(1-\lambda)f(\mathbf{y})
Key implication
H(x) ≻ 0 everywhere → f strictly convex → unique global minimum → GD converges to it. Convexity transforms optimization from hard search to guaranteed computation.
07 / 11
LA 101
M08 · L02
Quadratic Forms
Closed-Form Minimum
Quadratic Minimum
\mathbf{x}^* = -\tfrac{1}{2}A^{-1}\mathbf{b},\quad 2A\mathbf{x}+\mathbf{b}=\mathbf{0}
Applications
Least squares (A = XᵀX), Wiener filter (A = R_xx), ridge regression (A = XᵀX + λI). All solve a positive definite linear system via Cholesky. CG iterates minimize the quadratic in Krylov subspaces.
08 / 11
LA 101
M08 · L02
Constrained Optimization
Lagrange Multipliers
The Lagrangian
\mathcal{L}(\mathbf{x},\boldsymbol{\lambda}) = f(\mathbf{x}) - \boldsymbol{\lambda}^T\mathbf{g}(\mathbf{x})
At the constrained optimum, ∇f = λᵀ∇g (gradients are parallel). KKT conditions: stationarity + feasibility. Eigenvalue problems are constrained optimization — max xᵀAx s.t. ‖x‖ = 1 → top eigenvector.
10 / 11
LA 101
M08 · L02
Module 8 · Lesson 2 Complete
Optimization Unified
GD: first-order, linear convergence. Newton: second-order, quadratic convergence. Convexity: PD Hessian guarantees global optimum. Quadratics: closed-form via linear system. Lagrange: constraints via gradient alignment.
Module 8: Vector Calculus Connections
Gradients · Optimization · ML Applications
11 / 11