Home / LA 101 / Module 8 / Lesson 1
Stories Mode

Gradients and Jacobians

The gradient generalizes the derivative to multiple dimensions, the Jacobian extends it to vector-valued functions, and the Hessian captures second-order curvature — together they power all of modern optimization.

~20 min read M8 · L1 Intermediate

What Is a Gradient?

In single-variable calculus, the derivative f′(x) measures how much f changes per unit change in x. When f depends on multiple variables — say f(x₁, x₂, …, xₙ) — we need a richer object to capture all the directions in which f can change. That object is the gradient, denoted ∇f (read "nabla f" or "del f").

The gradient of a scalar-valued function f : ℝⁿ → ℝ at a point x is the vector of all its partial derivatives:

The Gradient
\nabla f(\mathbf{x}) = \begin{pmatrix} \frac{\partial f}{\partial x_1} \\ \frac{\partial f}{\partial x_2} \\ \vdots \\ \frac{\partial f}{\partial x_n} \end{pmatrix} \in \mathbb{R}^n
The gradient ∇f(x) ∈ ℝⁿ is the column vector of partial derivatives. Each component ∂f/∂xᵢ measures the instantaneous rate of change of f when only xᵢ varies. The gradient lives in the same space as the input x, not the output.

The gradient is always a column vector of the same dimension as the input x, even though f is scalar-valued. This dimensional convention is critical: when we write the chain rule or Newton's method in matrix form, the gradient must be a vector so that matrix-vector products are well-defined.

Partial derivatives are computed by differentiating with respect to one variable while holding all others constant. For example, if f(x₁, x₂) = x₁² + 3x₁x₂ + x₂³, then ∂f/∂x₁ = 2x₁ + 3x₂ and ∂f/∂x₂ = 3x₁ + 3x₂², so ∇f = [2x₁ + 3x₂, 3x₁ + 3x₂²]ᵀ.

The Gradient as Direction of Steepest Ascent

The deepest geometric insight about the gradient is its relationship to the directional derivative. Given a unit vector v ∈ ℝⁿ (‖v‖ = 1), the directional derivative of f in direction v at point x is defined as:

Directional Derivative
D_{\mathbf{v}} f(\mathbf{x}) = \lim_{h \to 0} \frac{f(\mathbf{x} + h\mathbf{v}) - f(\mathbf{x})}{h} = \nabla f(\mathbf{x}) \cdot \mathbf{v} = \nabla f(\mathbf{x})^T \mathbf{v}
The directional derivative D_v f(x) is the dot product of the gradient with the unit direction v. By the Cauchy-Schwarz inequality, |D_v f| ≤ ‖∇f‖, with equality when v is parallel to ∇f. This means ∇f points in the direction of greatest increase of f.

By the Cauchy-Schwarz inequality, the dot product ∇f(x) · v is maximized (over all unit vectors v) when v = ∇f(x) / ‖∇f(x)‖ — that is, when v points in the same direction as the gradient. The maximum directional derivative equals ‖∇f(x)‖. Conversely, f decreases most steeply in the direction −∇f(x).

This is the foundational principle behind gradient descent: to minimize f, repeatedly step in the direction of −∇f. Each step moves in the locally steepest downhill direction. The magnitude ‖∇f(x)‖ tells you how steep the surface is at x — a large gradient means you are on a steep slope, a small gradient means you are near flat terrain (possibly a minimum, maximum, or saddle point).

At a local minimum or maximum of a smooth function, ∇f(x) = 0 (the zero vector). This is the first-order necessary condition for optimality — the "flat" condition. Finding critical points (where the gradient vanishes) is the starting point of all optimization theory.

The Jacobian Matrix

When the function f maps ℝⁿ to ℝᵐ (not just to a single real number), each output component fᵢ has its own gradient. Stacking these gradients as rows produces the Jacobian matrix J, the fundamental linear approximation of a vector-valued function near a point.

The Jacobian Matrix
J(\mathbf{x}) = \frac{\partial \mathbf{f}}{\partial \mathbf{x}} = \begin{pmatrix} \frac{\partial f_1}{\partial x_1} & \cdots & \frac{\partial f_1}{\partial x_n} \\ \vdots & \ddots & \vdots \\ \frac{\partial f_m}{\partial x_1} & \cdots & \frac{\partial f_m}{\partial x_n} \end{pmatrix} \in \mathbb{R}^{m \times n}
The m × n Jacobian matrix J(x) of f : ℝⁿ → ℝᵐ has entry Jᵢⱼ = ∂fᵢ/∂xⱼ. Row i of J is the gradient of output component fᵢ. Column j of J shows how all outputs respond to perturbations in input xⱼ. For m = 1, J reduces to ∇f ᵀ (a row vector).

The Jacobian is the best linear approximation to f near x: f(x + δ) ≈ f(x) + J(x)δ for small perturbations δ. This linearization underlies Newton's method, sensitivity analysis, and the implicit function theorem. If you think of f as a machine transforming inputs to outputs, the Jacobian tells you exactly how small input changes propagate through the machine to produce output changes.

The determinant of the Jacobian (when m = n) measures how the function locally stretches or shrinks volumes. A Jacobian determinant of 2 means the function locally doubles all volumes; a determinant of −1 means it preserves volume but reverses orientation. In probability, the Jacobian determinant appears in the change-of-variables formula for multivariate distributions.

Worked Example: Polar Coordinates

Consider the transformation from polar to Cartesian coordinates: f(r, θ) = (r cos θ, r sin θ). The Jacobian is:

Polar Coordinates Jacobian
J(r,\theta) = \begin{pmatrix} \cos\theta & -r\sin\theta \\ \sin\theta & r\cos\theta \end{pmatrix}, \quad \det(J) = r
The Jacobian of the polar-to-Cartesian map has determinant det(J) = r cos²θ + r sin²θ = r. This is why the area element in polar coordinates is r dr dθ — the Jacobian determinant r converts between the two coordinate systems.

The Chain Rule in Matrix Form

One of the most powerful uses of the Jacobian is expressing the multivariable chain rule as a simple matrix product. Suppose g : ℝᵖ → ℝⁿ and f : ℝⁿ → ℝᵐ. The composition h = f ∘ g maps ℝᵖ → ℝᵐ. The chain rule states:

Matrix Chain Rule
J_h(\mathbf{t}) = J_f(\mathbf{g}(\mathbf{t})) \cdot J_g(\mathbf{t}), \quad \underbrace{(m \times p)}_{J_h} = \underbrace{(m \times n)}_{J_f} \cdot \underbrace{(n \times p)}_{J_g}
The Jacobian of a composition is the product of Jacobians. J_h(t) is m × p, J_f(g(t)) is m × n, and J_g(t) is n × p. Dimensions must be compatible for the matrix product to work. For scalar f and vector g, this reduces to ∇h(t) = J_g(t)ᵀ ∇f(g(t)).

This matrix chain rule is exactly what happens in backpropagation in neural networks. A deep network is a composition of many functions: f = fₗ ∘ fₗ₋₁ ∘ ⋯ ∘ f₁. The gradient of the loss with respect to the parameters of layer k is computed by multiplying Jacobians from the output layer back to layer k — that is, by repeated application of the matrix chain rule. Understanding backpropagation as iterated Jacobian multiplication is the key insight behind efficient automatic differentiation.

The chain rule also tells us something important about conditioning: if the Jacobians in a composition have large singular values, the gradients can explode; if they have small singular values, gradients can vanish. This is the mathematical origin of the vanishing/exploding gradient problem in deep networks — a linear algebra phenomenon hidden inside the chain rule.

The Hessian Matrix

Just as the Jacobian generalizes the first derivative, the Hessian matrix H(x) generalizes the second derivative to functions of multiple variables. For a twice-differentiable scalar function f : ℝⁿ → ℝ, the Hessian is the n × n matrix of all second-order partial derivatives:

The Hessian Matrix
H(\mathbf{x}) = \nabla^2 f(\mathbf{x}) = \begin{pmatrix} \frac{\partial^2 f}{\partial x_1^2} & \cdots & \frac{\partial^2 f}{\partial x_1 \partial x_n} \\ \vdots & \ddots & \vdots \\ \frac{\partial^2 f}{\partial x_n \partial x_1} & \cdots & \frac{\partial^2 f}{\partial x_n^2} \end{pmatrix} = H(\mathbf{x})^T
The Hessian H(x) is the n × n matrix with entry Hᵢⱼ = ∂²f/∂xᵢ∂xⱼ. By Schwarz's theorem (continuity of second partials), H is always symmetric: Hᵢⱼ = Hⱼᵢ. The Hessian is the Jacobian of the gradient: H = J(∇f).

The Hessian is always a symmetric matrix (when partial derivatives are continuous), so all the tools of Module 5 (eigendecomposition of symmetric matrices, positive definiteness) apply directly. The eigenvalues of H encode the curvature of f in the principal curvature directions:

The second-order Taylor expansion of f around a point x₀ is:

Second-Order Taylor Expansion
f(\mathbf{x}_0 + \boldsymbol{\delta}) \approx f(\mathbf{x}_0) + \nabla f(\mathbf{x}_0)^T \boldsymbol{\delta} + \tfrac{1}{2}\boldsymbol{\delta}^T H(\mathbf{x}_0) \boldsymbol{\delta}
The second-order Taylor expansion approximates f locally as a quadratic. The linear term ∇f(x₀)ᵀδ captures the slope; the quadratic term ½δᵀH(x₀)δ captures the curvature. This approximation is the foundation of Newton's method and quasi-Newton optimization algorithms.

The Hessian is the key object in Newton's method for optimization. At each iterate x_k, Newton's method minimizes the second-order Taylor expansion exactly: the update is δ = −H(x_k)⁻¹ ∇f(x_k), which requires solving a linear system with the Hessian. If H is positive definite (guaranteed near a local minimum), the update is well-defined and the method converges quadratically. This connects directly back to Cholesky decomposition (Module 7): solving the Newton system H δ = −∇f efficiently requires Cholesky factorization of the Hessian.

Gradient, Jacobian, and Hessian as a Hierarchy

It helps to see all three objects as a hierarchy of derivatives for scalar functions f : ℝⁿ → ℝ:

For vector-valued functions f : ℝⁿ → ℝᵐ, the first-order analog is the Jacobian J(x) ∈ ℝᵐˣⁿ (the first derivatives). There is no single "Hessian" for vector-valued f — instead, each output component fᵢ has its own Hessian Hᵢ(x). The full collection is sometimes called the "Jacobian of the Jacobian" or a third-order tensor.

Applications

Sensitivity Analysis

In engineering, you often need to know how sensitive the output of a system is to small changes in inputs. If f : ℝⁿ → ℝᵐ models a physical system, the Jacobian J(x) at the operating point x quantifies this sensitivity: a perturbation δx in the inputs produces an approximate output change J(x)δx. The singular values of J indicate which input directions have the most impact on outputs and which are nearly irrelevant.

Signal Processing and Filter Design

When fitting filter parameters by minimizing a loss function (e.g., mean squared error between desired and actual filter outputs), the gradient of the loss with respect to the filter coefficients is exactly the quantity gradient descent needs. In adaptive filtering algorithms like LMS (Least Mean Squares), the update rule is precisely: w_{k+1} = w_k − μ · ∇L(w_k), where L is the loss and ∇L is its gradient with respect to the filter weight vector w.

Automatic Differentiation

Modern deep learning frameworks (PyTorch, JAX, TensorFlow) implement automatic differentiation (autograd), which computes gradients exactly (not numerically) by applying the matrix chain rule through the computation graph. Forward mode AD computes the Jacobian-vector product J(x)v; reverse mode AD computes the vector-Jacobian product vᵀJ(x) (which gives ∇f when v = 1 for scalar f). The choice of mode depends on whether n (number of inputs) or m (number of outputs) is larger — reverse mode is preferred when m ≪ n, which is exactly the case for scalar loss functions in ML training.


Engineering Implication

In NumPy and SciPy, gradients are computed numerically via finite differences (numpy.gradient, scipy.optimize.approx_fprime). For analytic gradients of parametric models, PyTorch's .backward() and JAX's jax.grad() implement reverse-mode autodiff efficiently. The Jacobian is available via torch.autograd.functional.jacobian or jax.jacobian. The Hessian via torch.autograd.functional.hessian or jax.hessian — but computing the full Hessian costs O(n²) in memory and is typically avoided for large n; instead, Hessian-vector products (H v) are computed in O(n) using a second application of autodiff.

Key Takeaways

The gradient ∇f(x) ∈ ℝⁿ is the vector of first partial derivatives of a scalar function — it points in the direction of steepest ascent and vanishes at critical points. The Jacobian J(x) ∈ ℝᵐˣⁿ stacks the gradients of all m output components — it is the best linear approximation to a vector-valued function and gives the matrix chain rule its matrix form. The Hessian H(x) ∈ ℝⁿˣⁿ is the symmetric matrix of second partial derivatives — its eigenvalues encode curvature, its positive definiteness determines local minima, and it appears in Newton's method and the second-order Taylor expansion. Together these three objects form the mathematical foundation of all modern optimization in signal processing and machine learning.