ML 101
M06 · L01
Module 6

The Perceptron

The ancestor of all neural networks. One artificial neuron, a simple learning rule, and a theorem that launched a field — then a limitation that nearly ended it.

01 / 13
ML 101
M06 · L01
Inspiration

From Biology to Math

A biological neuron receives signals, integrates them, and fires if the total exceeds a threshold. McCulloch & Pitts formalized this in 1943. Rosenblatt made it trainable in 1957.

Input
Weighted sum
Output
Fire / no fire
02 / 13
ML 101
M06 · L01
The Model

Weighted Sum + Step Function

Multiply each input by its weight, sum them up, add a bias, and apply a hard threshold. Output 1 if the result is non-negative, 0 otherwise. That’s the entire computation.

  • Weights encode feature importance
  • Bias shifts the boundary off the origin
  • Step function makes the decision hard
03 / 13
ML 101
M06 · L01
Formulation

The Decision Rule

The perceptron computes a dot product of weights and inputs, adds the bias, then applies a threshold step. Geometrically, this defines a hyperplane in the input space.

Perceptron Output
\hat{y} = \theta(\mathbf{w}\cdot\mathbf{x}+b)
04 / 13
ML 101
M06 · L01
Geometry

A Line That Divides

The weight vector is perpendicular to the decision boundary. Class 1 lives on one side, class 0 on the other. Moving the bias shifts the line; rotating the weights tilts it.

Decision Boundary
w · x + b = 0  —  a hyperplane in input space
05 / 13
ML 101
M06 · L01
Learning Rule

Learn from Mistakes

For each training example, predict, check the error, and update. Correct predictions change nothing. Mistakes nudge the weights toward the correct answer by a small step η.

Weight Update
\mathbf{w} \leftarrow \mathbf{w} + \eta\,(y-\hat{y})\,\mathbf{x}
06 / 13
ML 101
M06 · L01
Convergence Theorem

Guaranteed to Converge

Rosenblatt proved: if the data is linearly separable, the perceptron will find a correct set of weights in a finite number of updates. The larger the margin, the fewer updates needed.

Condition
Lin. separable
Result
Finite steps
07 / 13
ML 101
M06 · L01
The Hard Limit

XOR: A Line Can’t Do It

XOR outputs 1 only when inputs differ. Plot (0,0), (0,1), (1,0), (1,1) — the 1s sit on opposite corners of a square. No single straight line can separate them. The perceptron fails, forever.

Truth Table
0 XOR 0 = 0  |  1 XOR 1 = 0
0 XOR 1 = 1  |  1 XOR 0 = 1
08 / 13
ML 101
M06 · L01
The Reckoning

Minsky, Papert & AI Winter

  • 1969: Perceptrons book proves single-layer limits rigorously
  • Widely read as “neural networks are a dead end”
  • Funding collapses — first AI Winter begins
  • 1986: Backpropagation revives the field
  • The fix was simple: add hidden layers
09 / 13
ML 101
M06 · L01
What It Can Do

Linear Separation Works

  • Any linearly separable binary classification
  • AND, OR, NAND, NOR gates — not XOR or XNOR
  • Equivalent to logistic regression with a hard threshold
  • Fast to train — one pass per example, no backprop
  • Cannot handle non-linear boundaries
10 / 13
ML 101
M06 · L01
Legacy

One Neuron, Infinite Descendants

  • Every neuron in a deep network is a perceptron with a smooth activation
  • Stack perceptrons in layers → represent any function
  • The update rule generalizes to gradient descent
  • Hidden layers solve XOR — and everything harder
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

The perceptron: weighted sum, step function, mistake-driven updates. Guaranteed to converge on linearly separable data — but powerless against XOR. The solution is hidden layers. That’s next.

13 / 13