LA 101
M08 · L02
Module 8: Vector Calculus Connections

Optimization with Linear Algebra

Gradient descent navigates loss landscapes step by step. Newton's method uses curvature for faster convergence. Lagrange multipliers enforce constraints. Together they make linear algebra the engine of all modern optimization.

01 / 11
LA 101
M08 · L02
Gradient Descent

Step in the Downhill Direction

−α∇f
Update direction
α
Step size (learning rate)
‖∇f‖→0
Convergence criterion
Convergence Rate
Convergence speed is governed by the condition number κ = λ_max/λ_min of the Hessian. High κ means slow convergence — the reason preconditioning matters.
02 / 11
LA 101
M08 · L02
The Update Rule

GD Equation

Gradient Descent Update
\mathbf{x}_{k+1} = \mathbf{x}_k - \alpha\,\nabla f(\mathbf{x}_k)
Step size intuition
α too large → overshoot and diverge. α too small → slow convergence. Optimal α ≈ 2/(λ_max + λ_min) for a quadratic. Line search finds the best α per iteration.
03 / 11
LA 101
M08 · L02
Loss Landscape

Convex vs Non-Convex

  • Convex: single bowl — any local min is global min
  • Non-convex: rugged — multiple local minima possible
  • Saddle points: ∇f = 0 but H indefinite — can trap GD
  • Flat plateaus: ‖∇f‖ ≈ 0 everywhere — very slow progress
  • SGD advantage: noise helps escape saddle points
04 / 11
LA 101
M08 · L02
Newton's Method

Use the Hessian

  • Second-order method: uses H(x_k) at each step
  • Quadratic convergence: correct digits double per step
  • Cost: O(n³) per step — infeasible for large n
  • Newton direction: δ = −H⁻¹∇f solves the local quadratic exactly
  • Quasi-Newton (L-BFGS): approximate H⁻¹ from gradient history
  • Safeguard: if H indefinite, add λI to force PD
05 / 11
LA 101
M08 · L02
Newton's Update Rule

Curvature-Aware Step

Newton Update
\mathbf{x}_{k+1} = \mathbf{x}_k - H(\mathbf{x}_k)^{-1}\nabla f(\mathbf{x}_k)
vs gradient descent
GD uses −α∇f (ignores curvature). Newton uses −H⁻¹∇f (accounts for curvature). In steep directions the step shrinks; in flat directions it grows. This is why Newton converges so fast near a minimum.
06 / 11
LA 101
M08 · L02
Convexity

PD Hessian → Global Min

Convexity Condition
f(\lambda\mathbf{x}+(1-\lambda)\mathbf{y}) \leq \lambda f(\mathbf{x})+(1-\lambda)f(\mathbf{y})
Key implication
H(x) ≻ 0 everywhere → f strictly convex → unique global minimum → GD converges to it. Convexity transforms optimization from hard search to guaranteed computation.
07 / 11
LA 101
M08 · L02
Quadratic Forms

Closed-Form Minimum

Quadratic Minimum
\mathbf{x}^* = -\tfrac{1}{2}A^{-1}\mathbf{b},\quad 2A\mathbf{x}+\mathbf{b}=\mathbf{0}
Applications
Least squares (A = XᵀX), Wiener filter (A = R_xx), ridge regression (A = XᵀX + λI). All solve a positive definite linear system via Cholesky. CG iterates minimize the quadratic in Krylov subspaces.
08 / 11
LA 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

09 / 11
LA 101
M08 · L02
Constrained Optimization

Lagrange Multipliers

The Lagrangian
\mathcal{L}(\mathbf{x},\boldsymbol{\lambda}) = f(\mathbf{x}) - \boldsymbol{\lambda}^T\mathbf{g}(\mathbf{x})

At the constrained optimum, ∇f = λᵀ∇g (gradients are parallel). KKT conditions: stationarity + feasibility. Eigenvalue problems are constrained optimization — max xᵀAx s.t. ‖x‖ = 1 → top eigenvector.

10 / 11
LA 101
M08 · L02
Module 8 · Lesson 2 Complete

Optimization Unified

GD: first-order, linear convergence. Newton: second-order, quadratic convergence. Convexity: PD Hessian guarantees global optimum. Quadratics: closed-form via linear system. Lagrange: constraints via gradient alignment.

Module 8: Vector Calculus Connections
Gradients · Optimization · ML Applications
11 / 11