ML 101
M02 · L02
Module 2

Logistic Regression

Linear regression predicts numbers. But what if the answer is yes or no? Spam or not spam? Tumor benign or malignant? Logistic regression transforms a line into a probability.

01 / 13
ML 101
M02 · L02
The Shift

From Numbers to Decisions

Linear regression outputs any number from −∞ to +∞. Classification needs a probability between 0 and 1. We need a function that squashes any input into that range.

Regression
−∞ to +∞
Classification
0 to 1
02 / 13
ML 101
M02 · L02
The Bridge

The Sigmoid Function

The sigmoid σ(z) maps any real number to (0, 1). Large positive → close to 1. Large negative → close to 0. Zero → exactly 0.5. It’s the S-curve that turns a linear model into a classifier.

Sigmoid Function
\sigma(z) = \frac{1}{1 + e^{-z}}
03 / 13
ML 101
M02 · L02
The Model

Predicting Probability

Logistic regression = linear model + sigmoid. Compute z = w·x + b, then pass through σ(z) to get the probability that the input belongs to class 1.

Logistic Model
P(y=1|\mathbf{x}) = \sigma(\mathbf{w}^T \mathbf{x} + b)
04 / 13
ML 101
M02 · L02
Classification

Decision Boundary

Where σ(z) = 0.5, i.e. z = 0, is the decision boundary. Points on one side are classified as 1, the other as 0. It’s a straight line (or hyperplane) in feature space.

Threshold
Predict class 1 if P(y=1|x) ≥ 0.5. The threshold can be adjusted for different precision/recall trade-offs.
05 / 13
ML 101
M02 · L02
The Right Loss

Why Not MSE?

MSE + sigmoid creates a non-convex loss landscape — full of local minima where gradient descent gets stuck. We need a loss function designed for probabilities: one that’s convex and heavily penalizes confident wrong predictions.

Key Insight
If you predict 0.99 for class 1 but the true label is 0, the penalty should be massive. MSE gives a mild 0.98. Cross-entropy gives ∞.
06 / 13
ML 101
M02 · L02
Loss Function

Cross-Entropy Loss

Binary cross-entropy measures the gap between predicted probabilities and true labels. It’s convex (one global minimum) and punishes confident mistakes harshly.

Binary Cross-Entropy
J = -\frac{1}{n}\sum_{i=1}^{n}\left[y_i \log(\hat{y}_i) + (1-y_i)\log(1-\hat{y}_i)\right]
07 / 13
ML 101
M02 · L02
Training

Gradient Descent

Same algorithm as linear regression — compute the gradient, step downhill. The gradient has an elegant form: the error (ŷ − y) times the input x. No closed-form solution exists, but gradient descent always converges since the loss is convex.

Update Rule
w ← w − α · Xᵀ (σ(Xw) − y) / n
08 / 13
ML 101
M02 · L02
Beyond Binary

Multi-Class Softmax

More than two classes? Softmax generalizes the sigmoid: it converts K raw scores into a probability distribution that sums to 1. Each class gets a probability.

Softmax Function
P(y=k|\mathbf{x}) = \frac{e^{z_k}}{\sum_{j=1}^{K} e^{z_j}}
09 / 13
ML 101
M02 · L02
Strategy

One-vs-Rest Classification

An alternative to softmax: train K separate binary classifiers, each answering “is it class k or not?” At prediction time, pick the class with the highest confidence.

Softmax
1 model, K outputs
One-vs-Rest
K models, 1 output
10 / 13
ML 101
M02 · L02
Prevention

Regularized Classification

Just like linear regression, logistic regression can overfit — especially with many features. Add L1 or L2 penalties to the cross-entropy loss. The parameter C = 1/λ controls the trade-off in most libraries.

  • L2 — smooth decision boundaries
  • L1 — sparse weights, feature selection
  • C large — less regularization
  • C small — more regularization
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What you learned

Logistic regression applies sigmoid to a linear model to output probabilities. Cross-entropy loss replaces MSE. Softmax extends to multi-class. Same gradient descent, different loss function.

Next Lesson
13 / 13