ML 101
M06 · L02
Module 6

Multi-Layer Perceptrons

One neuron draws a line. Stack layers of neurons and you can draw anything. Hidden layers break the linearity barrier — and unlock the full power of neural networks.

01 / 13
ML 101
M06 · L02
Architecture

Hidden Layers

Between inputs and output sit one or more hidden layers. Each layer transforms the representation. Each transformation can learn curved, irregular boundaries that no single hyperplane can express.

Input
Raw features
Hidden
Abstractions
02 / 13
ML 101
M06 · L02
Key Insight

Why Activations Matter

Composing linear transformations gives another linear transformation. To gain expressive power, every hidden neuron wraps its weighted sum in a non-linear activation — without this, depth buys nothing.

  • No activation → collapses to linear model
  • With activation → arbitrary function approximation
  • Choice of activation shapes gradient flow
03 / 13
ML 101
M06 · L02
Activation Zoo

Sigmoid & Tanh

The classics. Sigmoid squashes to (0,1), tanh to (−1,1). Smooth and differentiable. But they saturate at extremes: gradients shrink to near-zero, stalling learning in deep networks.

Sigmoid
\sigma(z)=\frac{1}{1+e^{-z}}
04 / 13
ML 101
M06 · L02
Modern Standard

ReLU — Simple Wins

max(0, z). Pass positives unchanged, kill negatives. Gradient is 1 for positive inputs — no saturation on the active side. Deep networks train dramatically faster. Today’s default for hidden layers.

Risk: Dead Neurons
If a neuron always gets negative input, its gradient is 0 forever — it never updates
05 / 13
ML 101
M06 · L02
ReLU Variants

Leaky, GELU, and Beyond

  • Leaky ReLU — small slope α for negatives; prevents dead neurons
  • Parametric ReLU — α is learned during training
  • GELU — Gaussian-gated; smooth; standard in transformers & LLMs
  • Swish — z · σ(z); self-gated; used in EfficientNet
06 / 13
ML 101
M06 · L02
Universal Approximation

One Layer to Rule Them All?

A single hidden layer of sufficient width can approximate any continuous function. Cybenko proved this in 1989. But “sufficient width” can be exponential — depth is far more efficient in practice.

Proven by
Cybenko 1989
Caveat
Exp. width
07 / 13
ML 101
M06 · L02
Computation

Forward Propagation

Layer by layer: multiply weights by the previous activations, add bias, apply the non-linearity. The output of each layer feeds the next. The final layer’s output is the prediction.

Layer l
\mathbf{z}^{[l]}=W^{[l]}\mathbf{a}^{[l-1]}+\mathbf{b}^{[l]},\quad\mathbf{a}^{[l]}=f\bigl(\mathbf{z}^{[l]}\bigr)
08 / 13
ML 101
M06 · L02
Output Layer

Task-Specific Activations

  • Sigmoid — binary classification, output in (0,1)
  • Softmax — multi-class, probabilities sum to 1
  • None (linear) — regression, raw predicted value
  • The choice is not optional — it must match the loss function
09 / 13
ML 101
M06 · L02
Design Choices

Depth vs. Width

  • Width — more neurons per layer, more patterns per step
  • Depth — more layers, hierarchical composition of features
  • Deeper networks: exponential expressiveness, polynomial parameters
  • Deeper networks: harder to train (vanishing gradients)
  • Modern practice: moderate depth + skip connections
10 / 13
ML 101
M06 · L02
Capacity & Overfitting

More Depth, More Risk

Each added layer increases capacity. Too much capacity and the network memorizes training noise instead of generalizing. More data or regularization is the fix — not fewer layers. The tools come in the next module.

Rule of Thumb
Overfit first, then regularize — confirms the model has enough capacity
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

Hidden layers + non-linear activations = universal function approximation. ReLU dominates hidden layers. Forward pass is layer-by-layer matrix multiplies. The output activation matches the task. Depth earns compositional power — at the cost of harder training.

Next Lesson
13 / 13