ML 101
M06 · L03
Module 6

Backpropagation

The network made a mistake. Now what? Backpropagation answers the hardest question in deep learning: which weights are to blame, and by exactly how much should each one change?

01 / 13
ML 101
M06 · L03
The Core Problem

Credit Assignment

Millions of weights, one error signal. A weight in layer 1 influences the output only through every layer above it. How do you fairly attribute blame across dozens of indirections?

Answer
Chain rule
Method
Backprop
02 / 13
ML 101
M06 · L03
The Foundation

The Chain Rule

A network is a nested composition of functions. The gradient of the loss with respect to any weight is a product of local derivatives along the path from that weight to the loss. That product is computed by the chain rule.

Chain Rule
\frac{\partial L}{\partial z} = \frac{\partial L}{\partial a} \cdot f'(z)
03 / 13
ML 101
M06 · L03
Pass 1

Forward Pass: Cache Everything

The forward pass computes the prediction — and stores every intermediate value. Pre-activations z[l] and activations a[l] at every layer are cached in memory. The backward pass will need them to compute local derivatives.

Memory Cost
All cached activations can consume gigabytes on deep networks — the price of efficiency
04 / 13
ML 101
M06 · L03
Output Layer

The First Delta

Start at the output. The error signal δ[L] is the gradient of the loss times the derivative of the output activation. For softmax + cross-entropy, this simplifies to just â − y — predicted minus true.

Output delta
\boldsymbol{\delta}^{[L]} = \nabla_{\mathbf{a}}L \odot f'\!\bigl(\mathbf{z}^{[L]}\bigr)
05 / 13
ML 101
M06 · L03
Pass 2: Backward

Propagating the Signal

Project the next layer’s delta back through the transposed weight matrix, then multiply element-wise by the local activation derivative. This produces δ[l] — the core recurrence that carries blame backward through every layer.

Hidden delta
\boldsymbol{\delta}^{[l]} = \bigl(W^{[l+1]\top}\boldsymbol{\delta}^{[l+1]}\bigr) \odot f'\!\bigl(\mathbf{z}^{[l]}\bigr)
06 / 13
ML 101
M06 · L03
Weight Updates

Gradient in Hand

Once δ[l] is known, the weight gradient is immediate: the outer product of δ[l] and the previous layer’s activations. The bias gradient is just δ[l] itself. Hand these to the optimizer and it updates the weights.

Weight gradient
\frac{\partial L}{\partial W^{[l]}} = \boldsymbol{\delta}^{[l]}\,\mathbf{a}^{[l-1]\top}
07 / 13
ML 101
M06 · L03
The Algorithm

The Training Loop

  • Sample a mini-batch of B examples
  • Forward pass: compute predictions, cache all z and a
  • Compute loss on batch predictions vs. labels
  • Backward pass: compute all deltas and gradients
  • Optimizer updates every weight — repeat
08 / 13
ML 101
M06 · L03
Pathology 1

Vanishing Gradients

Each layer multiplies the gradient by a number < 1. Sigmoid’s max derivative is 0.25. After 20 layers: 0.2520 ≈ 10−12. Early layers stop learning entirely — the network loses its hierarchical power.

Fix
ReLU activations, skip connections, batch normalization, LSTM gates
09 / 13
ML 101
M06 · L03
Pathology 2

Exploding Gradients

The opposite disaster: gradients grow exponentially layer by layer. Updates become enormous. Training diverges, NaN values appear. Common in recurrent networks with long sequences.

Fix
Gradient clipping — if ‖g‖ > threshold, scale g down so ‖g‖ = threshold
10 / 13
ML 101
Knowledge Check

Check whatstuck

Four questions on backpropagation — credit assignment, the forward cache, the output delta, and vanishing gradients.

Question 1 of 0
Score 0/0

11 / 13
ML 101
M06 · L03
Weight Initialization

Starting Right

  • All zeros — symmetry problem: all neurons identical, never diverge
  • Too large — activations saturate before training starts
  • Xavier — scale by 1/√nin; for sigmoid/tanh
  • He — scale by √(2/nin); for ReLU networks
12 / 13
ML 101
Summary
Recap

What You Learned

Backprop = chain rule applied backward. Forward pass caches activations; backward pass computes deltas layer by layer. Output delta is â−y for softmax+CE. Hidden delta projects through WT, then multiplies by f′(z). Gradients are outer products. Watch for vanishing/exploding gradients and initialize weights wisely.

Module Complete
Module 7: Regularization →
13 / 13