Reading
Stories Mode

Multi-Layer Perceptrons

~20 min read Lesson 2 of 3 in Module 6

Breaking the Linearity Barrier

The perceptron proved that a single artificial neuron could learn linearly separable patterns. But the XOR problem and countless real-world tasks demanded something more: the ability to draw curved, irregular, or intricate decision boundaries. The answer was hiding in plain sight — stack the perceptrons.

A multi-layer perceptron (MLP) adds one or more layers of neurons between the input and output. These hidden layers transform the raw features into new representations before the final classification or regression step. Each hidden layer can detect increasingly abstract patterns — edges from pixels, then shapes from edges, then objects from shapes. Depth earns expressive power.

The key insight: when you compose multiple linear transformations you still get a linear transformation. To gain true non-linearity, each neuron applies an activation function after its weighted sum. Without activation functions, a 100-layer network collapses to a single matrix multiplication and is no more powerful than a perceptron.

Activation Functions

An activation function wraps the linear weighted sum in a non-linear transformation. Different activation functions trade off between smoothness, gradient behavior, and computational cost. Choosing the right one for the right layer is part craft, part theory.

Sigmoid

The sigmoid squashes any real number into (0, 1), making it natural for binary output probabilities. It defined early neural networks. Its downfall: the gradient approaches zero for very large or very small inputs, causing the vanishing gradient problem during training. Rarely used in hidden layers today; still common in output layers for binary classification.

Sigmoid & Tanh
\sigma(z) = \frac{1}{1+e^{-z}} \qquad \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}}
Sigmoid maps inputs to (0,1). Tanh maps inputs to (−1, 1) and is zero-centered, which helps gradient flow. Both saturate at extremes — gradients become tiny when |z| is large, slowing learning in deep networks.

The Rectified Linear Unit (ReLU) changed everything. Defined simply as max(0, z), it passes positive values unchanged and kills negatives. This sounds almost too simple — but ReLU trains dramatically faster than sigmoid or tanh in deep networks. Its gradient is either 0 (for negative inputs) or 1 (for positive), which avoids saturation on the positive side and allows gradients to flow freely through many layers.

ReLU Variants
\text{ReLU}(z) = \max(0,z) \qquad \text{LeakyReLU}(z) = \max(\alpha z, z)
ReLU is the dominant activation for hidden layers. Leaky ReLU allows a small negative slope α (typically 0.01) for negative inputs, preventing “dead neurons” that never activate. GELU, used in transformers and modern LLMs, smoothly gates inputs using the Gaussian CDF and achieves better empirical performance on large language tasks.

Dead neurons are a real hazard with ReLU. If a neuron’s weights are updated such that its pre-activation output is always negative, it will output 0 for every input forever — its gradient is 0 so it never receives an update. Leaky ReLU, Parametric ReLU, and ELU all address this by keeping a small gradient alive on the negative side.

Network Architecture: Depth vs. Width

Every MLP is defined by the number of layers, the number of neurons per layer, and the activation functions. The input layer has one neuron per feature. The output layer has one neuron per class (for classification) or one neuron (for scalar regression). The hidden layers are the designer’s domain.

Width refers to the number of neurons in a layer. A wider layer can represent more patterns in a single transformation, but each additional neuron adds parameters. Depth refers to the number of hidden layers. Deeper networks can represent hierarchical abstractions with fewer total parameters than very wide shallow networks — but they are harder to train because gradients must flow through more layers.

Universal Approximation Theorem

A neural network with a single hidden layer of sufficient width and a non-linear activation function can approximate any continuous function on a compact domain to arbitrary accuracy. This result, proved by Cybenko (1989) for sigmoid and generalized by Hornik (1991) to other activations, establishes that MLPs are universal function approximators.

The caveat: “sufficient width” can be exponentially large. In practice, depth — not unlimited width — is the key to efficient approximation of the structured functions that appear in real problems. Two hidden layers of moderate width typically outperform one huge hidden layer.

Forward Propagation

Forward propagation is the process of computing the network’s output from an input. It proceeds layer by layer: the output of one layer becomes the input of the next. Each layer performs a linear transformation (matrix multiply plus bias) followed by a non-linear activation.

Layer Computation
\mathbf{z}^{[l]} = W^{[l]}\mathbf{a}^{[l-1]} + \mathbf{b}^{[l]} \qquad \mathbf{a}^{[l]} = f\!\left(\mathbf{z}^{[l]}\right)
For layer l, the pre-activation z is the matrix product of the weight matrix W&lsup;[l]&rsup; and the previous layer’s output a&lsup;[l−1]&rsup;, plus the bias vector b&lsup;[l]&rsup;. The activation a&lsup;[l]&rsup; applies the non-linearity element-wise. The final layer’s output is the network’s prediction.

In matrix form, processing an entire batch of examples simultaneously is straightforward — each column of the activation matrix is one example. This batched computation maps cleanly to GPU hardware, where thousands of floating-point operations run in parallel. It is one reason why deep learning became practical: GPUs turned forward propagation from a serial loop into a massive matrix multiply.

The output layer’s activation depends on the task. For binary classification, a sigmoid output gives a probability in (0, 1). For multi-class classification, a softmax converts the raw scores into a probability distribution over classes. For regression, no activation at all — the raw linear output is the prediction.

Softmax Output
\text{softmax}(z_k) = \frac{e^{z_k}}{\sum_{j=1}^{K} e^{z_j}}
The softmax function converts a vector of raw scores z into probabilities. Each entry is the exponential of its score divided by the sum of all exponentials. All outputs are positive and sum to one, making them interpretable as class probabilities. Softmax is numerically stabilized by subtracting max(z) before exponentiation.

Depth, Capacity, and Overfitting

Adding layers increases the capacity of the network — its ability to model complex functions. But capacity is a double-edged sword. A network with too much capacity relative to the training data will memorize the training examples rather than generalizing to new ones. This is overfitting: the network learns the noise in the data rather than the signal.

Deeper networks do not inherently overfit more than shallow ones — what matters is the ratio of parameters to informative training examples. A 10-layer network trained on 10 million labeled images may generalize beautifully; the same network trained on 100 examples will likely overfit badly. The practical tools for managing capacity — dropout, weight decay, early stopping — come in the next module.

Representational Hierarchy

The power of depth lies in composing representations. Layer 1 detects local patterns (edges, phonemes, n-grams). Layer 2 assembles those into larger structures (shapes, syllables, phrases). Layer 3 builds higher abstractions. This compositional reuse means deep networks can represent exponentially complex patterns using polynomially many neurons — a provable advantage over shallow networks for certain function classes.

Key Takeaways
  • Hidden layers let the network learn non-linear transformations of the input, breaking the linearity barrier that limits a single perceptron.
  • Activation functions are essential: without them, stacking linear layers produces only linear transformations. ReLU is the dominant choice for hidden layers today.
  • Sigmoid saturates and causes vanishing gradients; tanh is zero-centered but still saturates; ReLU trains fast but can produce dead neurons; Leaky ReLU and GELU address this.
  • The Universal Approximation Theorem guarantees that a single hidden layer of sufficient width can approximate any continuous function — but depth is more efficient in practice.
  • Forward propagation computes outputs layer by layer: z = Wa + b, then a = f(z). The whole computation maps to matrix multiplications, enabling GPU parallelism.
  • The output layer uses sigmoid for binary classification, softmax for multi-class, and no activation for regression.
  • Depth enables compositional representations and exponential expressiveness with polynomial parameters — a core reason deep learning dominates vision, speech, and language.
Previous The Perceptron Overview Next Backpropagation