Training Neural Networks
Backpropagation hands you the gradients. Now comes the craft: choosing the right loss, the right optimizer, the right learning rate, and the right regularization to make those gradients count.
What Are We Minimizing?
The loss function is the single number training reduces. Its choice must match the task: MSE for regression, cross-entropy for classification. Using the wrong loss creates subtle gradient problems that sabotage learning.
Mean Squared Error
MSE penalizes large errors quadratically. For classification tasks it creates a saturation problem — cross-entropy avoids this by producing a gradient always proportional to the prediction error.
Stochastic Gradient Descent
The simplest optimizer: subtract a fraction η of the gradient from each weight. Effective but slow on ill-conditioned loss surfaces — it oscillates across steep ravines instead of moving along them.
Adam & AdamW
Adam tracks running averages of gradients (momentum) and squared gradients (variance). Large, consistent gradients get a smaller step; small, noisy ones get larger. AdamW decouples weight decay so regularization works as intended.
Schedules & Warmup
A fixed learning rate is rarely optimal. Start large to converge fast, finish small to fine-tune. Cosine annealing smoothly follows a half-cosine curve. Warmup linearly ramps the rate for the first few hundred steps — critical for Adam.
Big vs. Small
- Large batches — fast GPU utilization, accurate gradients, but risk sharp minima that generalize poorly
- Small batches — noisier gradients that escape sharp minima and find flat, generalizable ones
- Linear scaling rule — multiply learning rate by k when multiplying batch size by k
Dropout: Noise as Defense
Randomly zero each neuron with probability p during training. Scaled up by 1/(1−p) to preserve expected activation. Prevents co-adaptation. At test time, all neurons are active — the network behaves like an average of exponentially many sub-networks.
Batch Normalization
Normalize each layer’s pre-activations across the mini-batch to zero mean and unit variance. Then apply learnable scale γ and shift β. Eliminates internal covariate shift, allows higher learning rates, acts as mild regularization.
The Default Recipe
- Loss: Cross-entropy (classification) or MSE (regression)
- Optimizer: AdamW, β1=0.9, β2=0.999, weight decay 0.01–0.1
- LR: Cosine annealing with linear warmup (5–10% of steps)
- Batch: 32–256; scale LR proportionally
- Regularization: Dropout + batch norm + weight decay
Systematic Ablation
The recipe is a starting point, not a formula. Vary one component at a time while holding others fixed. Larger datasets tolerate less regularization; larger models often need more. The scientific method applied to hyperparameters.
What You Learned
Loss must match the task. SGD is the foundation; Adam adapts per-parameter rates; AdamW fixes weight decay. Cosine annealing + warmup improves final performance. Small batches generalize better. Dropout and batch norm regularize and stabilize. Ablate systematically to understand each piece.