Logistic Regression
Linear regression predicts numbers. But what if the answer is yes or no? Spam or not spam? Tumor benign or malignant? Logistic regression transforms a line into a probability.
From Numbers to Decisions
Linear regression outputs any number from −∞ to +∞. Classification needs a probability between 0 and 1. We need a function that squashes any input into that range.
The Sigmoid Function
The sigmoid σ(z) maps any real number to (0, 1). Large positive → close to 1. Large negative → close to 0. Zero → exactly 0.5. It’s the S-curve that turns a linear model into a classifier.
Predicting Probability
Logistic regression = linear model + sigmoid. Compute z = w·x + b, then pass through σ(z) to get the probability that the input belongs to class 1.
Decision Boundary
Where σ(z) = 0.5, i.e. z = 0, is the decision boundary. Points on one side are classified as 1, the other as 0. It’s a straight line (or hyperplane) in feature space.
Why Not MSE?
MSE + sigmoid creates a non-convex loss landscape — full of local minima where gradient descent gets stuck. We need a loss function designed for probabilities: one that’s convex and heavily penalizes confident wrong predictions.
Cross-Entropy Loss
Binary cross-entropy measures the gap between predicted probabilities and true labels. It’s convex (one global minimum) and punishes confident mistakes harshly.
Gradient Descent
Same algorithm as linear regression — compute the gradient, step downhill. The gradient has an elegant form: the error (ŷ − y) times the input x. No closed-form solution exists, but gradient descent always converges since the loss is convex.
Multi-Class Softmax
More than two classes? Softmax generalizes the sigmoid: it converts K raw scores into a probability distribution that sums to 1. Each class gets a probability.
One-vs-Rest Classification
An alternative to softmax: train K separate binary classifiers, each answering “is it class k or not?” At prediction time, pick the class with the highest confidence.
Regularized Classification
Just like linear regression, logistic regression can overfit — especially with many features. Add L1 or L2 penalties to the cross-entropy loss. The parameter C = 1/λ controls the trade-off in most libraries.
- L2 — smooth decision boundaries
- L1 — sparse weights, feature selection
- C large — less regularization
- C small — more regularization
What you learned
Logistic regression applies sigmoid to a linear model to output probabilities. Cross-entropy loss replaces MSE. Softmax extends to multi-class. Same gradient descent, different loss function.