What Is a Gradient?
In single-variable calculus, the derivative f′(x) measures how much f changes per unit change in x. When f depends on multiple variables — say f(x₁, x₂, …, xₙ) — we need a richer object to capture all the directions in which f can change. That object is the gradient, denoted ∇f (read "nabla f" or "del f").
The gradient of a scalar-valued function f : ℝⁿ → ℝ at a point x is the vector of all its partial derivatives:
The gradient is always a column vector of the same dimension as the input x, even though f is scalar-valued. This dimensional convention is critical: when we write the chain rule or Newton's method in matrix form, the gradient must be a vector so that matrix-vector products are well-defined.
Partial derivatives are computed by differentiating with respect to one variable while holding all others constant. For example, if f(x₁, x₂) = x₁² + 3x₁x₂ + x₂³, then ∂f/∂x₁ = 2x₁ + 3x₂ and ∂f/∂x₂ = 3x₁ + 3x₂², so ∇f = [2x₁ + 3x₂, 3x₁ + 3x₂²]ᵀ.
The Gradient as Direction of Steepest Ascent
The deepest geometric insight about the gradient is its relationship to the directional derivative. Given a unit vector v ∈ ℝⁿ (‖v‖ = 1), the directional derivative of f in direction v at point x is defined as:
By the Cauchy-Schwarz inequality, the dot product ∇f(x) · v is maximized (over all unit vectors v) when v = ∇f(x) / ‖∇f(x)‖ — that is, when v points in the same direction as the gradient. The maximum directional derivative equals ‖∇f(x)‖. Conversely, f decreases most steeply in the direction −∇f(x).
This is the foundational principle behind gradient descent: to minimize f, repeatedly step in the direction of −∇f. Each step moves in the locally steepest downhill direction. The magnitude ‖∇f(x)‖ tells you how steep the surface is at x — a large gradient means you are on a steep slope, a small gradient means you are near flat terrain (possibly a minimum, maximum, or saddle point).
At a local minimum or maximum of a smooth function, ∇f(x) = 0 (the zero vector). This is the first-order necessary condition for optimality — the "flat" condition. Finding critical points (where the gradient vanishes) is the starting point of all optimization theory.
The Jacobian Matrix
When the function f maps ℝⁿ to ℝᵐ (not just to a single real number), each output component fᵢ has its own gradient. Stacking these gradients as rows produces the Jacobian matrix J, the fundamental linear approximation of a vector-valued function near a point.
The Jacobian is the best linear approximation to f near x: f(x + δ) ≈ f(x) + J(x)δ for small perturbations δ. This linearization underlies Newton's method, sensitivity analysis, and the implicit function theorem. If you think of f as a machine transforming inputs to outputs, the Jacobian tells you exactly how small input changes propagate through the machine to produce output changes.
The determinant of the Jacobian (when m = n) measures how the function locally stretches or shrinks volumes. A Jacobian determinant of 2 means the function locally doubles all volumes; a determinant of −1 means it preserves volume but reverses orientation. In probability, the Jacobian determinant appears in the change-of-variables formula for multivariate distributions.
Worked Example: Polar Coordinates
Consider the transformation from polar to Cartesian coordinates: f(r, θ) = (r cos θ, r sin θ). The Jacobian is:
The Chain Rule in Matrix Form
One of the most powerful uses of the Jacobian is expressing the multivariable chain rule as a simple matrix product. Suppose g : ℝᵖ → ℝⁿ and f : ℝⁿ → ℝᵐ. The composition h = f ∘ g maps ℝᵖ → ℝᵐ. The chain rule states:
This matrix chain rule is exactly what happens in backpropagation in neural networks. A deep network is a composition of many functions: f = fₗ ∘ fₗ₋₁ ∘ ⋯ ∘ f₁. The gradient of the loss with respect to the parameters of layer k is computed by multiplying Jacobians from the output layer back to layer k — that is, by repeated application of the matrix chain rule. Understanding backpropagation as iterated Jacobian multiplication is the key insight behind efficient automatic differentiation.
The chain rule also tells us something important about conditioning: if the Jacobians in a composition have large singular values, the gradients can explode; if they have small singular values, gradients can vanish. This is the mathematical origin of the vanishing/exploding gradient problem in deep networks — a linear algebra phenomenon hidden inside the chain rule.
The Hessian Matrix
Just as the Jacobian generalizes the first derivative, the Hessian matrix H(x) generalizes the second derivative to functions of multiple variables. For a twice-differentiable scalar function f : ℝⁿ → ℝ, the Hessian is the n × n matrix of all second-order partial derivatives:
The Hessian is always a symmetric matrix (when partial derivatives are continuous), so all the tools of Module 5 (eigendecomposition of symmetric matrices, positive definiteness) apply directly. The eigenvalues of H encode the curvature of f in the principal curvature directions:
- All eigenvalues positive: H is positive definite, x is a strict local minimum of f.
- All eigenvalues negative: H is negative definite, x is a strict local maximum of f.
- Mixed signs: H is indefinite, x is a saddle point.
- Any zero eigenvalue: the second-order test is inconclusive; higher-order terms determine the nature of the critical point.
The second-order Taylor expansion of f around a point x₀ is:
The Hessian is the key object in Newton's method for optimization. At each iterate x_k, Newton's method minimizes the second-order Taylor expansion exactly: the update is δ = −H(x_k)⁻¹ ∇f(x_k), which requires solving a linear system with the Hessian. If H is positive definite (guaranteed near a local minimum), the update is well-defined and the method converges quadratically. This connects directly back to Cholesky decomposition (Module 7): solving the Newton system H δ = −∇f efficiently requires Cholesky factorization of the Hessian.
Gradient, Jacobian, and Hessian as a Hierarchy
It helps to see all three objects as a hierarchy of derivatives for scalar functions f : ℝⁿ → ℝ:
- Zeroth order: f(x) ∈ ℝ — the value itself.
- First order: ∇f(x) ∈ ℝⁿ — the gradient, a vector of first partials. Points in the direction of steepest ascent.
- Second order: H(x) ∈ ℝⁿˣⁿ — the Hessian, a matrix of second partials. Encodes curvature. Always symmetric. Related to ∇f by H = J(∇f).
For vector-valued functions f : ℝⁿ → ℝᵐ, the first-order analog is the Jacobian J(x) ∈ ℝᵐˣⁿ (the first derivatives). There is no single "Hessian" for vector-valued f — instead, each output component fᵢ has its own Hessian Hᵢ(x). The full collection is sometimes called the "Jacobian of the Jacobian" or a third-order tensor.
Applications
Sensitivity Analysis
In engineering, you often need to know how sensitive the output of a system is to small changes in inputs. If f : ℝⁿ → ℝᵐ models a physical system, the Jacobian J(x) at the operating point x quantifies this sensitivity: a perturbation δx in the inputs produces an approximate output change J(x)δx. The singular values of J indicate which input directions have the most impact on outputs and which are nearly irrelevant.
Signal Processing and Filter Design
When fitting filter parameters by minimizing a loss function (e.g., mean squared error between desired and actual filter outputs), the gradient of the loss with respect to the filter coefficients is exactly the quantity gradient descent needs. In adaptive filtering algorithms like LMS (Least Mean Squares), the update rule is precisely: w_{k+1} = w_k − μ · ∇L(w_k), where L is the loss and ∇L is its gradient with respect to the filter weight vector w.
Automatic Differentiation
Modern deep learning frameworks (PyTorch, JAX, TensorFlow) implement automatic differentiation (autograd), which computes gradients exactly (not numerically) by applying the matrix chain rule through the computation graph. Forward mode AD computes the Jacobian-vector product J(x)v; reverse mode AD computes the vector-Jacobian product vᵀJ(x) (which gives ∇f when v = 1 for scalar f). The choice of mode depends on whether n (number of inputs) or m (number of outputs) is larger — reverse mode is preferred when m ≪ n, which is exactly the case for scalar loss functions in ML training.
In NumPy and SciPy, gradients are computed numerically via finite differences (numpy.gradient, scipy.optimize.approx_fprime). For analytic gradients of parametric models, PyTorch's .backward() and JAX's jax.grad() implement reverse-mode autodiff efficiently. The Jacobian is available via torch.autograd.functional.jacobian or jax.jacobian. The Hessian via torch.autograd.functional.hessian or jax.hessian — but computing the full Hessian costs O(n²) in memory and is typically avoided for large n; instead, Hessian-vector products (H v) are computed in O(n) using a second application of autodiff.
The gradient ∇f(x) ∈ ℝⁿ is the vector of first partial derivatives of a scalar function — it points in the direction of steepest ascent and vanishes at critical points. The Jacobian J(x) ∈ ℝᵐˣⁿ stacks the gradients of all m output components — it is the best linear approximation to a vector-valued function and gives the matrix chain rule its matrix form. The Hessian H(x) ∈ ℝⁿˣⁿ is the symmetric matrix of second partial derivatives — its eigenvalues encode curvature, its positive definiteness determines local minima, and it appears in Newton's method and the second-order Taylor expansion. Together these three objects form the mathematical foundation of all modern optimization in signal processing and machine learning.