ML 101
M6 · Lab
Hands-on lab
Watch one gradient flow back

Step a tiny network through a forward pass, a backward pass and one update in the sandbox, and predict — before you read it off the screen — the output ŷ and loss, one weight gradient ∂L/∂w, and the weight's new value after a single gradient-descent step.

One gradient-descent step
w \leftarrow w - \eta\,\dfrac{\partial L}{\partial w}

Backpropagation is just the chain rule bookkeeping which stored value each gradient reuses. The forward pass computes values and stores them; the backward pass reuses them to compute ∂L/∂w for every weight; then one gradient-descent step nudges each weight downhill, w ← w − η ∂L/∂w. Doing all three on one concrete network, with numbers you can check, is what turns backprop from a formula into something you can see.

1 / 8
ML 101
M6 · Lab
Set it up
One tiny network

Open the sandbox and set the base knobs below. Then, in the editable parameter grid, type this hand-built network exactly — the clean weights make every number below checkable by hand. Every step then walks the same stepper and grid.

Set these values
Depth -> 1 Width -> 2 Activation -> sigmoid Loss -> MSE Data -> one row x = [1, 0] y = 1 W[1] = [[1,0],[0,0]] b[1] = [0,0] W[2] = [[1,1]] b[2] = [0]

The output unit is always a sigmoid, so ŷ stays a probability. The stepper shows the forward values, the grid prints ∂L/∂w beside each weight, and the update step prints each new weight — every number you predict is right there.

2 / 8
ML 101
M6 · Lab
Step 1 of 4
Forward pass: the output

Step to Output ŷ in the stepper. With x = [1, 0] the hidden pre-activations are z[1] = [1, 0], so a[1] = [σ(1), σ(0)] = [0.7311, 0.5], then z[2] = 1·0.7311 + 1·0.5 = 1.2311. Predict ŷ = σ(1.2311), then read it.

Expected

The output reads ŷ = 0.7740 — a single number in (0, 1). That is the whole forward pass: layer by layer, values in, one probability out.

3 / 8
ML 101
M6 · Lab
Step 2 of 4
Forward pass: the loss

Step to Loss L. The target is y = 1 and the loss is MSE, written L = (ŷ − y)² for this one row. Predict L from ŷ = 0.7740, then read it.

Expected

The loss reads L = 0.0511 — that is (0.7740 − 1)². One scalar summarising how far this prediction landed from the target; the backward pass exists to make it smaller.

4 / 8
ML 101
M6 · Lab
Step 3 of 4
Backward pass: one gradient

Now look at the editable grid, at the output weight W[2]11. The backward pass reuses the stored a[1]1 = 0.7311: the gradient is ∂L/∂W[2]11 = δ[2] · a[1]1, with the output delta δ[2] = 2(ŷ−y)·ŷ(1−ŷ) = −0.0791. Predict ∂L/∂W[2]11, then read it beside the weight.

Expected

The grid prints ∂L/∂W[2]11 = -0.0578. Negative means nudging this weight up lowers the loss — and notice it reused a value the forward pass already stored, rather than recomputing it. That reuse is the whole point of backpropagation.

5 / 8
ML 101
M6 · Lab
Step 4 of 4
The update: one step downhill

Set the learning rate to η = 1 (the Learning rate slider at 0 on its log scale) and step to the update. Gradient descent moves the weight against its gradient: W[2]11 ← W[2]11 − η ∂L/∂W[2]11 = 1 − 1·(−0.0578). Predict the new W[2]11, then read it.

Expected

The update step prints W[2]11 ← 1.0578. Because the gradient was negative, the weight went up, and one such step for every weight at once is exactly what training does — thousands of times.

6 / 8
ML 101
M6 · Lab
Your turn
Open the sandbox

Everything above is waiting in the sandbox. Type any weights into the grid, step through forward and backward and watch each gradient reuse a stored value, then apply updates and watch the loss fall. Switch the activation to ReLU or stack hidden layers to see gradients vanish.

7 / 8
ML 101
M6 · Lab
Wrap-up
What you did
  • Ran a forward pass to the output ŷ = 0.7740
  • Read the MSE loss L = 0.0511 for target y = 1
  • Watched the backward pass produce one gradient, ∂L/∂W[2]11 = -0.0578, by reusing a stored value
  • Took one gradient-descent step at η = 1: W[2]11 ← 1.0578
  • Every value you predicted is the demo's own forward and backward arithmetic, not a picture
8 / 8