From Regression to Classification
Linear regression predicts a continuous number — a price, a temperature, a score. But many real-world problems ask a different question: which category? Is this email spam or not? Is this tumor benign or malignant? Will the customer churn or stay? These are classification problems, and they need a model that outputs probabilities, not unbounded numbers.
The Sigmoid Function
The key insight: take the output of a linear model (which ranges from −∞ to +∞) and squash it into the range (0, 1) using the sigmoid function. The result is interpreted as the probability that the input belongs to class 1.
The logistic regression model is simply: compute z = w·x + b (a linear combination), then apply σ(z) to get the predicted probability.
The Decision Boundary
To make a prediction, we apply a threshold: if P(y=1|x) ≥ 0.5, predict class 1; otherwise predict class 0. The set of points where σ(z) = 0.5 (equivalently, z = 0) forms the decision boundary — a straight line (or hyperplane) in feature space that separates the two classes.
Cross-Entropy Loss
Why not use MSE for classification? Because MSE combined with the sigmoid creates a non-convex loss landscape riddled with local minima. Instead, we use binary cross-entropy (also called log loss), which is convex and specifically designed for probabilistic outputs.
If the true label is 1 and you predict 0.99, the loss is −log(0.99) ≈ 0.01 — tiny. But predict 0.01 and the loss is −log(0.01) ≈ 4.6 — massive. Cross-entropy destroys overconfident wrong answers, exactly what we want for classification.
Multi-Class Classification
Binary classification covers two classes. For K > 2 classes, two approaches exist. One-vs-Rest trains K separate binary classifiers, each asking “is it class k or not?” Softmax regression is more elegant: it generalizes the sigmoid to K classes, producing a probability distribution over all classes that sums to 1.
Regularization in Classification
Logistic regression can overfit too — especially with many features or linearly separable data (weights grow to ±∞ trying to perfectly separate classes). Add L1 or L2 penalties to the cross-entropy loss, just as with linear regression. Most libraries use the parameter C = 1/λ: large C means less regularization, small C means more.
- Logistic regression applies the sigmoid function to a linear model, converting unbounded outputs into probabilities between 0 and 1.
- The decision boundary is where P(y=1|x) = 0.5, forming a hyperplane in feature space.
- Cross-entropy loss replaces MSE — it’s convex and heavily penalizes confident wrong predictions.
- Softmax extends logistic regression to multi-class problems, outputting a probability distribution over all classes.
- Regularization (L1/L2) prevents overfitting by penalizing large weights, controlled by the parameter C = 1/λ.