Backpropagation
The network made a mistake. Now what? Backpropagation answers the hardest question in deep learning: which weights are to blame, and by exactly how much should each one change?
Credit Assignment
Millions of weights, one error signal. A weight in layer 1 influences the output only through every layer above it. How do you fairly attribute blame across dozens of indirections?
The Chain Rule
A network is a nested composition of functions. The gradient of the loss with respect to any weight is a product of local derivatives along the path from that weight to the loss. That product is computed by the chain rule.
Forward Pass: Cache Everything
The forward pass computes the prediction — and stores every intermediate value. Pre-activations z[l] and activations a[l] at every layer are cached in memory. The backward pass will need them to compute local derivatives.
The First Delta
Start at the output. The error signal δ[L] is the gradient of the loss times the derivative of the output activation. For softmax + cross-entropy, this simplifies to just â − y — predicted minus true.
Propagating the Signal
Project the next layer’s delta back through the transposed weight matrix, then multiply element-wise by the local activation derivative. This produces δ[l] — the core recurrence that carries blame backward through every layer.
Gradient in Hand
Once δ[l] is known, the weight gradient is immediate: the outer product of δ[l] and the previous layer’s activations. The bias gradient is just δ[l] itself. Hand these to the optimizer and it updates the weights.
The Training Loop
- Sample a mini-batch of B examples
- Forward pass: compute predictions, cache all z and a
- Compute loss on batch predictions vs. labels
- Backward pass: compute all deltas and gradients
- Optimizer updates every weight — repeat
Vanishing Gradients
Each layer multiplies the gradient by a number < 1. Sigmoid’s max derivative is 0.25. After 20 layers: 0.2520 ≈ 10−12. Early layers stop learning entirely — the network loses its hierarchical power.
Exploding Gradients
The opposite disaster: gradients grow exponentially layer by layer. Updates become enormous. Training diverges, NaN values appear. Common in recurrent networks with long sequences.
Starting Right
- All zeros — symmetry problem: all neurons identical, never diverge
- Too large — activations saturate before training starts
- Xavier — scale by 1/√nin; for sigmoid/tanh
- He — scale by √(2/nin); for ReLU networks
What You Learned
Backprop = chain rule applied backward. Forward pass caches activations; backward pass computes deltas layer by layer. Output delta is â−y for softmax+CE. Hidden delta projects through WT, then multiplies by f′(z). Gradients are outer products. Watch for vanishing/exploding gradients and initialize weights wisely.