LA 101
M08 · L01
Module 8: Vector Calculus Connections
Gradients and Jacobians
The gradient generalizes the derivative to multiple dimensions. The Jacobian extends it to vector-valued functions. Together they power all of modern optimization and automatic differentiation.
01 / 11
LA 101
M08 · L01
The Gradient
Vector of Partial Derivatives
∇f ∈ ℝⁿ
Same dimension as input
∂f/∂xᵢ
i-th component
∇f = 0
Critical point condition
Interpretation
Each component ∂f/∂xᵢ is the rate of change of f when only xᵢ varies, holding all others fixed. The gradient stacks all these rates into a single vector.
02 / 11
LA 101
M08 · L01
The Gradient Formula
∇f Points Uphill
Gradient Definition
\nabla f(\mathbf{x}) = \begin{pmatrix}\partial f/\partial x_1\\ \vdots \\ \partial f/\partial x_n\end{pmatrix} \in \mathbb{R}^n
Steepest ascent
D_v f = ∇f · v is maximized when v ∥ ∇f. Gradient descent steps in direction −∇f — the steepest downhill direction.
03 / 11
LA 101
M08 · L01
Directional Derivative
How Fast in Any Direction?
Directional Derivative
D_{\mathbf{v}}f(\mathbf{x}) = \nabla f(\mathbf{x})^T \mathbf{v}, \quad \max_{\|\mathbf{v}\|=1} D_{\mathbf{v}}f = \|\nabla f\|
The gradient is the unique vector that encodes all directional derivatives: knowing ∇f at a point tells you D_v f in every direction v via a single dot product.
04 / 11
LA 101
M08 · L01
The Jacobian Matrix
Gradient for Vector Functions
- f : ℝⁿ → ℝᵐ — m outputs, n inputs
- J ∈ ℝᵐˣⁿ — m × n matrix, Jᵢⱼ = ∂fᵢ/∂xⱼ
- Row i — gradient of output component fᵢ
- Column j — sensitivity of all outputs to input xⱼ
- f(x + δ) ≈ f(x) + Jδ — best linear approximation
- det(J) — local volume scaling factor
05 / 11
LA 101
M08 · L01
The Jacobian Formula
Matrix of Partials
Jacobian Matrix
J_{ij} = \frac{\partial f_i}{\partial x_j},\quad J \in \mathbb{R}^{m\times n},\quad \mathbf{f}(\mathbf{x}+\boldsymbol{\delta})\approx \mathbf{f}(\mathbf{x})+J\boldsymbol{\delta}
Polar coordinates example
For f(r,θ) = (r cosθ, r sinθ), det(J) = r. This explains the extra r in the polar area element dA = r dr dθ.
06 / 11
LA 101
M08 · L01
Matrix Chain Rule
Jacobians Multiply
Chain Rule for Compositions
J_h = J_f \cdot J_g,\quad \underbrace{(m{\times}p)}_{J_h} = \underbrace{(m{\times}n)}_{J_f}\cdot\underbrace{(n{\times}p)}_{J_g}
Backpropagation
Deep networks are compositions f = fₗ ∘ ⋯ ∘ f₁. Backprop is iterated Jacobian multiplication — from output to input. Vanishing/exploding gradients arise from small/large Jacobian singular values.
07 / 11
LA 101
M08 · L01
The Hessian Matrix
Second-Order Curvature
Second-Order Taylor Expansion
f(\mathbf{x}_0+\boldsymbol{\delta})\approx f(\mathbf{x}_0)+\nabla f^T\boldsymbol{\delta}+\tfrac{1}{2}\boldsymbol{\delta}^T H\boldsymbol{\delta}
Hessian H = ∇²f
Symmetric n × n matrix of second partials. Positive definite → local min. Negative definite → local max. Indefinite → saddle point. Used in Newton's method: δ = −H⁻¹∇f.
08 / 11
LA 101
M08 · L01
Applications
Where Gradients Appear
- Gradient descent: x_{k+1} = x_k − α∇f(x_k)
- Newton's method: x_{k+1} = x_k − H⁻¹∇f(x_k)
- Backpropagation: Jacobian chain rule through networks
- Adaptive filtering (LMS): weight update via ∇L
- Sensitivity analysis: Jδx estimates output changes
- Autodiff: forward mode = Jv, reverse mode = vᵀJ
10 / 11
LA 101
M08 · L01
Module 8 · Lesson 1 Complete
∇f, J, H: The Calculus Trio
Gradient ∇f points uphill. Jacobian J linearizes vector maps. Hessian H captures curvature. Together they are the mathematical backbone of optimization — from signal processing to deep learning.
Module 8: Vector Calculus Connections
Gradients · Optimization · ML Applications
11 / 11