Reading
Stories Mode

Logistic Regression

~14 min read Lesson 2 of Module 2

From Regression to Classification

Linear regression predicts a continuous number — a price, a temperature, a score. But many real-world problems ask a different question: which category? Is this email spam or not? Is this tumor benign or malignant? Will the customer churn or stay? These are classification problems, and they need a model that outputs probabilities, not unbounded numbers.

The Sigmoid Function

The key insight: take the output of a linear model (which ranges from −∞ to +∞) and squash it into the range (0, 1) using the sigmoid function. The result is interpreted as the probability that the input belongs to class 1.

Sigmoid Function
\sigma(z) = \frac{1}{1 + e^{-z}}
Maps any real number to (0, 1). Large positive inputs → near 1. Large negative inputs → near 0. At z = 0, the output is exactly 0.5.

The logistic regression model is simply: compute z = w·x + b (a linear combination), then apply σ(z) to get the predicted probability.

Logistic Model
P(y=1|\mathbf{x}) = \sigma(\mathbf{w}^T \mathbf{x} + b)
The probability that input x belongs to class 1, parameterized by weights w and bias b.

The Decision Boundary

To make a prediction, we apply a threshold: if P(y=1|x) ≥ 0.5, predict class 1; otherwise predict class 0. The set of points where σ(z) = 0.5 (equivalently, z = 0) forms the decision boundary — a straight line (or hyperplane) in feature space that separates the two classes.

Cross-Entropy Loss

Why not use MSE for classification? Because MSE combined with the sigmoid creates a non-convex loss landscape riddled with local minima. Instead, we use binary cross-entropy (also called log loss), which is convex and specifically designed for probabilistic outputs.

Binary Cross-Entropy
-\frac{1}{n}\sum_{i=1}^{n}\left[y_i \log(\hat{y}_i) + (1-y_i)\log(1-\hat{y}_i)\right]
When y=1, only −log(ŷ) matters. When y=0, only −log(1−ŷ) matters. Confident wrong predictions are penalized heavily.
Why Cross-Entropy Works

If the true label is 1 and you predict 0.99, the loss is −log(0.99) ≈ 0.01 — tiny. But predict 0.01 and the loss is −log(0.01) ≈ 4.6 — massive. Cross-entropy destroys overconfident wrong answers, exactly what we want for classification.

Multi-Class Classification

Binary classification covers two classes. For K > 2 classes, two approaches exist. One-vs-Rest trains K separate binary classifiers, each asking “is it class k or not?” Softmax regression is more elegant: it generalizes the sigmoid to K classes, producing a probability distribution over all classes that sums to 1.

Softmax Function
P(y=k|\mathbf{x}) = \frac{e^{z_k}}{\sum_{j=1}^{K} e^{z_j}}
Converts K raw scores (logits) into a probability distribution. Each class gets a probability; they all sum to 1.

Regularization in Classification

Logistic regression can overfit too — especially with many features or linearly separable data (weights grow to ±∞ trying to perfectly separate classes). Add L1 or L2 penalties to the cross-entropy loss, just as with linear regression. Most libraries use the parameter C = 1/λ: large C means less regularization, small C means more.

In the next lesson, we’ll dive deep into regularization — L1, L2, and Elastic Net — and understand why penalizing large weights is one of the most important techniques in all of machine learning.

Key Takeaways
  • Logistic regression applies the sigmoid function to a linear model, converting unbounded outputs into probabilities between 0 and 1.
  • The decision boundary is where P(y=1|x) = 0.5, forming a hyperplane in feature space.
  • Cross-entropy loss replaces MSE — it’s convex and heavily penalizes confident wrong predictions.
  • Softmax extends logistic regression to multi-class problems, outputting a probability distribution over all classes.
  • Regularization (L1/L2) prevents overfitting by penalizing large weights, controlled by the parameter C = 1/λ.
PreviousLinear Regression Module Overview Next LessonRegularization