The Perceptron
The ancestor of all neural networks. One artificial neuron, a simple learning rule, and a theorem that launched a field — then a limitation that nearly ended it.
From Biology to Math
A biological neuron receives signals, integrates them, and fires if the total exceeds a threshold. McCulloch & Pitts formalized this in 1943. Rosenblatt made it trainable in 1957.
Weighted Sum + Step Function
Multiply each input by its weight, sum them up, add a bias, and apply a hard threshold. Output 1 if the result is non-negative, 0 otherwise. That’s the entire computation.
- Weights encode feature importance
- Bias shifts the boundary off the origin
- Step function makes the decision hard
The Decision Rule
The perceptron computes a dot product of weights and inputs, adds the bias, then applies a threshold step. Geometrically, this defines a hyperplane in the input space.
A Line That Divides
The weight vector is perpendicular to the decision boundary. Class 1 lives on one side, class 0 on the other. Moving the bias shifts the line; rotating the weights tilts it.
Learn from Mistakes
For each training example, predict, check the error, and update. Correct predictions change nothing. Mistakes nudge the weights toward the correct answer by a small step η.
Guaranteed to Converge
Rosenblatt proved: if the data is linearly separable, the perceptron will find a correct set of weights in a finite number of updates. The larger the margin, the fewer updates needed.
XOR: A Line Can’t Do It
XOR outputs 1 only when inputs differ. Plot (0,0), (0,1), (1,0), (1,1) — the 1s sit on opposite corners of a square. No single straight line can separate them. The perceptron fails, forever.
0 XOR 1 = 1 | 1 XOR 0 = 1
Minsky, Papert & AI Winter
- 1969: Perceptrons book proves single-layer limits rigorously
- Widely read as “neural networks are a dead end”
- Funding collapses — first AI Winter begins
- 1986: Backpropagation revives the field
- The fix was simple: add hidden layers
Linear Separation Works
- Any linearly separable binary classification
- AND, OR, NAND, NOR gates — not XOR or XNOR
- Equivalent to logistic regression with a hard threshold
- Fast to train — one pass per example, no backprop
- Cannot handle non-linear boundaries
One Neuron, Infinite Descendants
- Every neuron in a deep network is a perceptron with a smooth activation
- Stack perceptrons in layers → represent any function
- The update rule generalizes to gradient descent
- Hidden layers solve XOR — and everything harder
What You Learned
The perceptron: weighted sum, step function, mistake-driven updates. Guaranteed to converge on linearly separable data — but powerless against XOR. The solution is hidden layers. That’s next.