From Gradients to a Working Model
Backpropagation gives you the gradients. Now what? The art of training neural networks lies in the choices surrounding those gradients: what loss function to minimize, which optimizer to use, how fast to update weights, how large each batch should be, and how to regularize the network so it generalizes to new data. These decisions have an outsized impact on whether a model converges at all, how fast it trains, and whether it overfits.
This lesson covers the practical machinery of training — the components you assemble every time you build and train a neural network in the real world.
Loss Functions: What Are We Minimizing?
The loss function is the single number that training tries to reduce. Its choice is not arbitrary — it must match the task and the output distribution you are modeling.
For regression (predicting a continuous value), the standard choice is Mean Squared Error (MSE). It penalizes large errors quadratically, making it sensitive to outliers but mathematically convenient: the gradient is simply the prediction error scaled by 2/n.
For classification, the dominant choice is cross-entropy loss (also called log loss). It measures the dissimilarity between the predicted probability distribution and the true label distribution. Combined with a softmax output layer, cross-entropy produces the elegant gradient δ[L] = â − y that we derived in the previous lesson.
MSE combined with sigmoid/softmax outputs creates a flat loss landscape when predictions are very wrong — the gradient nearly vanishes because the sigmoid is near saturation. Cross-entropy avoids this: its gradient is always proportional to the prediction error â − y, providing a strong signal even when the model is confidently wrong.
Optimizers: How to Apply the Gradients
The optimizer determines how the gradients are used to update the weights. The simplest optimizer is Stochastic Gradient Descent (SGD): subtract a fraction η (the learning rate) of the gradient from each weight. This is the same learning rate written α in the gradient-descent introduction (Lesson 2.1); η is simply the conventional symbol for it in deep learning.
Pure SGD is effective but slow on ill-conditioned loss surfaces (where the landscape curves steeply in some directions and gently in others). SGD with momentum addresses this by accumulating a velocity vector that dampens oscillations and accelerates movement along consistent gradient directions — like a ball rolling downhill that picks up speed.
Modern deep learning overwhelmingly uses Adam (Adaptive Moment Estimation), which combines momentum with adaptive per-parameter learning rates. Adam maintains two running averages: a first moment estimate (mean of gradients) and a second moment estimate (mean of squared gradients). Parameters that receive large, consistent gradients get a smaller effective learning rate; parameters with small, noisy gradients get a larger one.
AdamW is a widely used variant that decouples weight decay from the gradient update. In standard Adam, L2 regularization (weight decay) is absorbed into the gradient — which means it also gets adapted by the second moment, reducing its effect. AdamW applies weight decay directly to the weights before the gradient step, restoring the regularization’s intended strength. Most state-of-the-art models use AdamW.
Learning Rate Schedules
A fixed learning rate is rarely optimal. Early in training a large learning rate accelerates convergence; late in training a small rate allows fine-tuning to the minimum. Learning rate schedules vary the rate over the course of training.
Step decay reduces the learning rate by a fixed factor (e.g., ×0.1) every fixed number of epochs. Simple and interpretable, but requires manual tuning of when to drop.
Cosine annealing smoothly decreases the learning rate following a cosine curve from its initial value to near zero. It avoids the abrupt drops of step decay and has been shown to produce better final performance. Many modern recipes use cosine annealing with periodic restarts (SGDR), where the learning rate is reset and decayed again cyclically.
Warmup is a technique used at the start of training: the learning rate is linearly increased from a small value to the target rate over the first few hundred or thousand steps. Warmup is particularly important for Adam-family optimizers, which have poorly calibrated second moment estimates at initialization. Without warmup, early large steps can destabilize training or push the model into a bad region of the loss landscape.
Batch Size Effects
The batch size — the number of training examples processed in one forward/backward pass — is one of the most consequential hyperparameters, though its effect is subtle.
Large batches provide accurate gradient estimates (low noise), allow efficient GPU parallelism, and train faster in wall-clock time per epoch. However, large-batch training is well-documented to converge to sharp minima — points in the loss landscape with steep surrounding walls. Sharp minima generalize poorly: a small distribution shift can cause the model to fall off the minimum. There is also a practical ceiling: extremely large batches provide diminishing returns since the gradient estimate is already near-exact.
Small batches introduce gradient noise that acts as implicit regularization, helping escape sharp minima and find flatter ones that generalize better. The downside is that each parameter update is noisier, requiring more updates to converge — and small batches under-utilize GPU parallelism.
When increasing batch size by a factor k, multiply the learning rate by k as well. This preserves the expected weight update magnitude. For example, going from batch size 256 to 2048 (8×) should be paired with an 8× increase in learning rate. This rule holds well in practice but breaks down at very large batch sizes, where warmup becomes especially important.
Dropout: Regularization by Noise
Dropout is a remarkably simple and effective regularization technique. During each training step, every neuron in a dropout layer is independently set to zero with probability p (the dropout rate, typically 0.1–0.5). The surviving neurons’ outputs are scaled up by 1/(1−p) to preserve the expected activation magnitude.
Dropout prevents neurons from co-adapting — developing complex joint dependencies where one neuron’s output only makes sense in combination with another’s. Because any neuron might be dropped at any time, each neuron must learn to be useful independently. The result is an implicit ensemble: at test time (when dropout is disabled and all neurons are active), the network behaves like an average of exponentially many sub-networks.
Dropout is only active during training. During evaluation and inference, all neurons are present and their outputs are multiplied by (1−p) — or equivalently, outputs are scaled up by 1/(1−p) during training (inverted dropout, the standard implementation). Failing to disable dropout at inference time is a common bug that degrades model performance.
Batch Normalization: Stabilizing the Training Distribution
As a network trains, the distribution of inputs to each layer shifts as the weights of the preceding layers change. This phenomenon — called internal covariate shift — forces later layers to continually adapt to a moving target, slowing training. Batch normalization (BN) addresses this by normalizing each layer’s pre-activation across the mini-batch to have zero mean and unit variance, then applying learnable scale (γ) and shift (β) parameters.
Batch normalization has several beneficial side effects beyond stabilizing distributions. It allows higher learning rates (training is less sensitive to initialization), acts as a mild regularizer (the noise from batch statistics introduces variability during training), and reduces the importance of careful weight initialization. Most modern architectures include batch normalization or its variants (layer normalization, group normalization) as standard components.
The original paper applies BN before the activation function (pre-activation BN). Many modern implementations — particularly ResNets — use pre-activation ResNets where BN appears before both the activation and the skip connection, which improves gradient flow. In practice, either placement usually works; consistency within a model matters more than the choice itself.
A Complete Training Recipe
These components combine into a standard training setup for a modern deep learning model. While exact choices vary by architecture and task, a reasonable default recipe looks like:
Loss: Cross-entropy for classification, MSE or Huber loss for regression.
Optimizer: AdamW with β1 = 0.9, β2 = 0.999, ε = 10−8, weight decay = 0.01–0.1.
Learning rate: Cosine annealing with linear warmup over 5–10% of total steps. Initial rate ~10−3 for Adam-family optimizers, ~10−1 for SGD.
Batch size: 32–256 for most tasks. Scale learning rate proportionally when increasing batch size.
Regularization: Dropout (p = 0.1–0.3) after dense layers. Batch normalization after linear/conv layers before activation. Weight decay in the optimizer.
Initialization: He for ReLU networks, Xavier for sigmoid/tanh.
This recipe is a starting point, not a formula. The right choices depend on dataset size (more data tolerates less regularization), model size (larger models often need stronger regularization), and task structure. A large fraction of deep learning practice involves systematic ablation — varying one component at a time while holding others fixed — to understand what each piece contributes.
- The loss function must match the task: MSE for regression, cross-entropy for classification. Cross-entropy avoids gradient saturation that MSE suffers with sigmoid outputs.
- SGD is the foundation; momentum accelerates it on ill-conditioned surfaces. Adam adapts per-parameter learning rates using first and second moment estimates. AdamW decouples weight decay, restoring its regularization effect.
- Learning rate schedules — step decay, cosine annealing, warmup — improve final performance by starting aggressive and finishing precise. Warmup is critical for Adam-family optimizers.
- Large batches train faster but risk sharp minima that generalize poorly. Small batches introduce useful noise. The linear scaling rule: multiply learning rate by k when multiplying batch size by k.
- Dropout randomly zeroes activations during training, preventing co-adaptation and acting as an implicit ensemble. It must be disabled at inference time.
- Batch normalization normalizes layer inputs across the mini-batch, stabilizing distributions, enabling higher learning rates, and acting as mild regularization. Learnable γ and β parameters restore representational capacity.
- A complete training recipe combines all these choices. Systematic ablation — varying one component at a time — is the standard method for understanding their individual contributions.