LA 101
M08 · L01
Module 8: Vector Calculus Connections

Gradients and Jacobians

The gradient generalizes the derivative to multiple dimensions. The Jacobian extends it to vector-valued functions. Together they power all of modern optimization and automatic differentiation.

01 / 11
LA 101
M08 · L01
The Gradient

Vector of Partial Derivatives

∇f ∈ ℝⁿ
Same dimension as input
∂f/∂xᵢ
i-th component
∇f = 0
Critical point condition
Interpretation
Each component ∂f/∂xᵢ is the rate of change of f when only xᵢ varies, holding all others fixed. The gradient stacks all these rates into a single vector.
02 / 11
LA 101
M08 · L01
The Gradient Formula

∇f Points Uphill

Gradient Definition
\nabla f(\mathbf{x}) = \begin{pmatrix}\partial f/\partial x_1\\ \vdots \\ \partial f/\partial x_n\end{pmatrix} \in \mathbb{R}^n
Steepest ascent
D_v f = ∇f · v is maximized when v ∥ ∇f. Gradient descent steps in direction −∇f — the steepest downhill direction.
03 / 11
LA 101
M08 · L01
Directional Derivative

How Fast in Any Direction?

Directional Derivative
D_{\mathbf{v}}f(\mathbf{x}) = \nabla f(\mathbf{x})^T \mathbf{v}, \quad \max_{\|\mathbf{v}\|=1} D_{\mathbf{v}}f = \|\nabla f\|

The gradient is the unique vector that encodes all directional derivatives: knowing ∇f at a point tells you D_v f in every direction v via a single dot product.

04 / 11
LA 101
M08 · L01
The Jacobian Matrix

Gradient for Vector Functions

  • f : ℝⁿ → ℝᵐ — m outputs, n inputs
  • J ∈ ℝᵐˣⁿ — m × n matrix, Jᵢⱼ = ∂fᵢ/∂xⱼ
  • Row i — gradient of output component fᵢ
  • Column j — sensitivity of all outputs to input xⱼ
  • f(x + δ) ≈ f(x) + Jδ — best linear approximation
  • det(J) — local volume scaling factor
05 / 11
LA 101
M08 · L01
The Jacobian Formula

Matrix of Partials

Jacobian Matrix
J_{ij} = \frac{\partial f_i}{\partial x_j},\quad J \in \mathbb{R}^{m\times n},\quad \mathbf{f}(\mathbf{x}+\boldsymbol{\delta})\approx \mathbf{f}(\mathbf{x})+J\boldsymbol{\delta}
Polar coordinates example
For f(r,θ) = (r cosθ, r sinθ), det(J) = r. This explains the extra r in the polar area element dA = r dr dθ.
06 / 11
LA 101
M08 · L01
Matrix Chain Rule

Jacobians Multiply

Chain Rule for Compositions
J_h = J_f \cdot J_g,\quad \underbrace{(m{\times}p)}_{J_h} = \underbrace{(m{\times}n)}_{J_f}\cdot\underbrace{(n{\times}p)}_{J_g}
Backpropagation
Deep networks are compositions f = fₗ ∘ ⋯ ∘ f₁. Backprop is iterated Jacobian multiplication — from output to input. Vanishing/exploding gradients arise from small/large Jacobian singular values.
07 / 11
LA 101
M08 · L01
The Hessian Matrix

Second-Order Curvature

Second-Order Taylor Expansion
f(\mathbf{x}_0+\boldsymbol{\delta})\approx f(\mathbf{x}_0)+\nabla f^T\boldsymbol{\delta}+\tfrac{1}{2}\boldsymbol{\delta}^T H\boldsymbol{\delta}
Hessian H = ∇²f
Symmetric n × n matrix of second partials. Positive definite → local min. Negative definite → local max. Indefinite → saddle point. Used in Newton's method: δ = −H⁻¹∇f.
08 / 11
LA 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

09 / 11
LA 101
M08 · L01
Applications

Where Gradients Appear

  • Gradient descent: x_{k+1} = x_k − α∇f(x_k)
  • Newton's method: x_{k+1} = x_k − H⁻¹∇f(x_k)
  • Backpropagation: Jacobian chain rule through networks
  • Adaptive filtering (LMS): weight update via ∇L
  • Sensitivity analysis: Jδx estimates output changes
  • Autodiff: forward mode = Jv, reverse mode = vᵀJ
10 / 11
LA 101
M08 · L01
Module 8 · Lesson 1 Complete

∇f, J, H: The Calculus Trio

Gradient ∇f points uphill. Jacobian J linearizes vector maps. Hessian H captures curvature. Together they are the mathematical backbone of optimization — from signal processing to deep learning.

Module 8: Vector Calculus Connections
Gradients · Optimization · ML Applications
11 / 11