Multi-Layer Perceptrons
One neuron draws a line. Stack layers of neurons and you can draw anything. Hidden layers break the linearity barrier — and unlock the full power of neural networks.
Hidden Layers
Between inputs and output sit one or more hidden layers. Each layer transforms the representation. Each transformation can learn curved, irregular boundaries that no single hyperplane can express.
Why Activations Matter
Composing linear transformations gives another linear transformation. To gain expressive power, every hidden neuron wraps its weighted sum in a non-linear activation — without this, depth buys nothing.
- No activation → collapses to linear model
- With activation → arbitrary function approximation
- Choice of activation shapes gradient flow
Sigmoid & Tanh
The classics. Sigmoid squashes to (0,1), tanh to (−1,1). Smooth and differentiable. But they saturate at extremes: gradients shrink to near-zero, stalling learning in deep networks.
ReLU — Simple Wins
max(0, z). Pass positives unchanged, kill negatives. Gradient is 1 for positive inputs — no saturation on the active side. Deep networks train dramatically faster. Today’s default for hidden layers.
Leaky, GELU, and Beyond
- Leaky ReLU — small slope α for negatives; prevents dead neurons
- Parametric ReLU — α is learned during training
- GELU — Gaussian-gated; smooth; standard in transformers & LLMs
- Swish — z · σ(z); self-gated; used in EfficientNet
One Layer to Rule Them All?
A single hidden layer of sufficient width can approximate any continuous function. Cybenko proved this in 1989. But “sufficient width” can be exponential — depth is far more efficient in practice.
Forward Propagation
Layer by layer: multiply weights by the previous activations, add bias, apply the non-linearity. The output of each layer feeds the next. The final layer’s output is the prediction.
Task-Specific Activations
- Sigmoid — binary classification, output in (0,1)
- Softmax — multi-class, probabilities sum to 1
- None (linear) — regression, raw predicted value
- The choice is not optional — it must match the loss function
Depth vs. Width
- Width — more neurons per layer, more patterns per step
- Depth — more layers, hierarchical composition of features
- Deeper networks: exponential expressiveness, polynomial parameters
- Deeper networks: harder to train (vanishing gradients)
- Modern practice: moderate depth + skip connections
More Depth, More Risk
Each added layer increases capacity. Too much capacity and the network memorizes training noise instead of generalizing. More data or regularization is the fix — not fewer layers. The tools come in the next module.
What You Learned
Hidden layers + non-linear activations = universal function approximation. ReLU dominates hidden layers. Forward pass is layer-by-layer matrix multiplies. The output activation matches the task. Depth earns compositional power — at the cost of harder training.