Pick a quadratic surface, set a learning rate η, and step in the cross-course optimizer demo, and predict — before the screen tells you — the gradient ∇f at a point, where one step x ← x − η∇f(x) lands, the exact learning rate at which plain descent stops converging, and what happens just past it.
Gradient descent is the workhorse of optimization, and on a quadratic f(x) = ½ x⊤Ax it is pure linear algebra. The gradient is ∇f = Ax; one step maps x to (I − ηA)x, so the whole run converges exactly when every eigenvalue of I − ηA is inside the unit circle — that is, when η < 2 / λmax. The eigenvalues of A set both the speed and the stability limit.
Open the demo (it lives in the ML 101 course, linked from this lab). Each step names the surface, the learning rate η and the start point to set; read the gradient arrows, the trajectory, the loss panel and the 2 / λmax threshold marker on the learning-rate track.
Surface = Bowl eta = 0.50 start = (-2.4, 2.0)
Read: grad f at a point x after one step 2 / lambda_max threshold
Watch the gradient arrows on the contour, the trajectory as you step, the loss read-out, and the threshold mark on the learning-rate slider — every number you predict is on screen.
Choose the Bowl surface, f(x,y) = ½(x2 + y2), and put the start at (−2.4, 2.0). Its gradient is ∇f = [x, y]. Predict the two components of ∇f at that point, then read the gradient arrow off the plot.
∇f(−2.4, 2.0) = [−2.4, 2.0]. On the bowl the gradient equals your position vector, so it points straight away from the minimum at the origin — descent moves you directly toward it.
Keep the Bowl and the start (−2.4, 2.0), and set the learning rate η = 0.50. Take exactly one step: x ← x − η∇f(x). Predict the x-coordinate the iterate lands on after that single step.
It lands at (−1.2, 1.0): −2.4 − 0.5·(−2.4) = −1.2 and 2.0 − 0.5·2.0 = 1.0. The loss drops from 4.88 to 1.22 — one step halves the distance to the minimum on this well-conditioned bowl.
Switch to the Ravine, f(x,y) = ½(30x2 + y2) — the same bowl stretched 30:1. Its Hessian A has eigenvalues λ = 30 and 1, so λmax = 30. Predict the exact learning rate 2 / λmax at which plain descent is on the edge of diverging, then read the threshold mark on the track.
The threshold is 2 / 30 = 0.0667. Below it every |1 − ηλ| < 1 and the loss falls each step; above it the steep direction (λ = 30) has |1 − ηλ| > 1 and the iterates blow up. The single steep eigenvalue sets the ceiling for the whole run.
Stay on the Ravine and push the learning rate to η = 0.08 — just above the 0.0667 threshold. Run the trajectory. Predict whether plain gradient descent converges or diverges.
It diverges. At η = 0.08 the steep direction has |1 − 0.08·30| = 1.4 > 1, so every step multiplies that component by 1.4 and the iterates run off to infinity — the demo flags the run as diverged. Crossing 2 / λmax by a hair is the difference between converging and blowing up.
Everything above is waiting in the demo. Drag the learning rate past the threshold mark and watch a converging run turn into a diverging one, switch surfaces to feel how the eigenvalue spread changes the picture, and step through the trajectory one gradient at a time.
- Read the gradient ∇f(−2.4, 2.0) = [−2.4, 2.0] off the bowl — on a bowl it equals your position
- Took one step x ← x − η∇f at η = 0.5 and landed at (−1.2, 1.0), loss 4.88 → 1.22
- Found the exact convergence threshold 2 / λmax = 0.0667 on the ravine
- Pushed η to 0.08 and watched plain descent diverge — the eigenvalues set the stability limit
- Every value you predicted is the demo's own gradient / step / eigenvalue arithmetic, not a drawn curve