ML 101
M06 · L04
Module 6

Training Neural Networks

Backpropagation hands you the gradients. Now comes the craft: choosing the right loss, the right optimizer, the right learning rate, and the right regularization to make those gradients count.

01 / 13
ML 101
M06 · L04
Loss Functions

What Are We Minimizing?

The loss function is the single number training reduces. Its choice must match the task: MSE for regression, cross-entropy for classification. Using the wrong loss creates subtle gradient problems that sabotage learning.

Regression
MSE / MAE
Classification
Cross-entropy
02 / 13
ML 101
M06 · L04
Regression Loss

Mean Squared Error

MSE penalizes large errors quadratically. For classification tasks it creates a saturation problem — cross-entropy avoids this by producing a gradient always proportional to the prediction error.

MSE & Cross-Entropy
\mathcal{L}_{\text{MSE}}=\tfrac{1}{n}\sum(\hat{y}_i-y_i)^2 \qquad \mathcal{L}_{\text{CE}}=-\log\hat{p}_{\text{true}}
03 / 13
ML 101
M06 · L04
Optimizers

Stochastic Gradient Descent

The simplest optimizer: subtract a fraction η of the gradient from each weight. Effective but slow on ill-conditioned loss surfaces — it oscillates across steep ravines instead of moving along them.

SGD Update
w \leftarrow w - \eta\,\nabla_w L(w)
04 / 13
ML 101
M06 · L04
Modern Optimizers

Adam & AdamW

Adam tracks running averages of gradients (momentum) and squared gradients (variance). Large, consistent gradients get a smaller step; small, noisy ones get larger. AdamW decouples weight decay so regularization works as intended.

Default settings
β1=0.9, β2=0.999, ε=10−8
05 / 13
ML 101
M06 · L04
Learning Rate

Schedules & Warmup

A fixed learning rate is rarely optimal. Start large to converge fast, finish small to fine-tune. Cosine annealing smoothly follows a half-cosine curve. Warmup linearly ramps the rate for the first few hundred steps — critical for Adam.

Cosine Annealing
\eta_t = \eta_{\min}+\tfrac{1}{2}(\eta_{\max}-\eta_{\min})\!\left(1+\cos\!\frac{t\pi}{T}\right)
06 / 13
ML 101
M06 · L04
Batch Size

Big vs. Small

  • Large batches — fast GPU utilization, accurate gradients, but risk sharp minima that generalize poorly
  • Small batches — noisier gradients that escape sharp minima and find flat, generalizable ones
  • Linear scaling rule — multiply learning rate by k when multiplying batch size by k
07 / 13
ML 101
M06 · L04
Regularization

Dropout: Noise as Defense

Randomly zero each neuron with probability p during training. Scaled up by 1/(1−p) to preserve expected activation. Prevents co-adaptation. At test time, all neurons are active — the network behaves like an average of exponentially many sub-networks.

Typical rate
p = 0.1–0.5
Inference
Disabled
08 / 13
ML 101
M06 · L04
Stabilization

Batch Normalization

Normalize each layer’s pre-activations across the mini-batch to zero mean and unit variance. Then apply learnable scale γ and shift β. Eliminates internal covariate shift, allows higher learning rates, acts as mild regularization.

Batch Norm
\hat{x}=\frac{x-\mu_B}{\sqrt{\sigma_B^2+\epsilon}}\quad y=\gamma\hat{x}+\beta
09 / 13
ML 101
M06 · L04
Putting It Together

The Default Recipe

  • Loss: Cross-entropy (classification) or MSE (regression)
  • Optimizer: AdamW, β1=0.9, β2=0.999, weight decay 0.01–0.1
  • LR: Cosine annealing with linear warmup (5–10% of steps)
  • Batch: 32–256; scale LR proportionally
  • Regularization: Dropout + batch norm + weight decay
10 / 13
ML 101
Knowledge Check

Check whatstuck

Four questions on training craft — choosing the loss, the SGD step, batch size, and dropout.

Question 1 of 0
Score 0/0

11 / 13
ML 101
M06 · L04
Tuning Practice

Systematic Ablation

The recipe is a starting point, not a formula. Vary one component at a time while holding others fixed. Larger datasets tolerate less regularization; larger models often need more. The scientific method applied to hyperparameters.

Rule of thumb
Change one thing at a time and measure the effect on validation loss
12 / 13
ML 101
Summary
Recap

What You Learned

Loss must match the task. SGD is the foundation; Adam adapts per-parameter rates; AdamW fixes weight decay. Cosine annealing + warmup improves final performance. Small batches generalize better. Dropout and batch norm regularize and stabilize. Ablate systematically to understand each piece.

Module Complete
Module 7: Regularization →
13 / 13