ML 101
M05 · L01
Module 5

Clustering — K-Means

No labels. No teacher. Just data and structure. Unsupervised learning finds hidden groupings — and K-Means is the algorithm that launched a thousand segmentations.

01 / 13
ML 101
M05 · L01
The Task

Grouping Without Labels

Clustering partitions data into groups so that within-group similarity is high and between-group similarity is low. No ground truth required — the structure emerges from the data itself.

Input
n points
Output
K groups
02 / 13
ML 101
M05 · L01
The Loop

Assign → Update → Repeat

  • Initialize — place K centroids (randomly or via K-Means++)
  • Assign — each point joins its nearest centroid
  • Update — recompute each centroid as the cluster mean
  • Check — stop if assignments did not change
  • Guaranteed to converge; converges to a local minimum
03 / 13
ML 101
M05 · L01
Objective

Minimizing Inertia

K-Means minimizes the total within-cluster sum of squared distances from each point to its centroid. Lower inertia = tighter, more compact clusters.

K-Means Objective
J = \displaystyle\sum_{k=1}^{K} \sum_{\mathbf{x} \in C_k} \|\mathbf{x} - \boldsymbol{\mu}_k\|^2
04 / 13
ML 101
M05 · L01
Hyperparameter

Choosing K

  • Elbow method — plot inertia vs. K; look for the bend
  • Silhouette score — measures cohesion vs. separation per point
  • Range −1 (wrong cluster) to +1 (perfect); maximize over K
  • Domain knowledge often most reliable guide
  • Neither metric is definitive — treat as heuristics
05 / 13
ML 101
M05 · L01
Quality Metric

Silhouette Score

a(i) = avg distance to same-cluster points. b(i) = avg distance to nearest other cluster. s(i) near +1 means well-clustered; near 0 means borderline.

Silhouette Formula
s(i) = \dfrac{b(i) - a(i)}{\max(a(i),\,b(i))}
06 / 13
ML 101
M05 · L01
Smarter Start

K-Means++

Random initialization risks poor local minima. K-Means++ seeds centroids by choosing each new one with probability proportional to its squared distance from the nearest existing centroid — spreading them out intelligently.

Guarantee
O(log K) approximation to optimal • Now the default in scikit-learn
07 / 13
ML 101
M05 · L01
Limitations

When K-Means Fails

  • Spherical assumption — can't find curved or elongated clusters
  • Equal sizes — large clusters get split; small ones merged
  • Outlier sensitivity — one extreme point pulls a centroid far away
  • K required upfront — must guess or search
  • Euclidean only — not suited for categorical or text data
08 / 13
ML 101
M05 · L01
At Scale

Mini-Batch K-Means

Standard K-Means loads all data per iteration. Mini-Batch K-Means updates centroids on a random subsample each step — like SGD for clustering. Dramatically faster on millions of points.

Speed
Much faster
Quality
Slightly ↑ inertia
09 / 13
ML 101
M05 · L01
Practical Note

Scale Before Clustering

K-Means uses Euclidean distance. A feature with large values dominates all distance calculations. Always standardize to zero mean, unit variance before running K-Means.

Rule of Thumb
StandardScaler before KMeans • Most common source of poor clustering
10 / 13
ML 101
M05 · L01
Real World

Where K-Means Shines

  • Customer segmentation — group users by behavior
  • Image compression — reduce colors to K representative values
  • Document clustering — topic discovery in large corpora
  • Anomaly detection — points far from any centroid are outliers
  • Feature engineering — cluster membership as a new feature
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

K-Means iterates assign→update to minimize inertia. K-Means++ seeds smartly. The elbow and silhouette help choose K. Mini-Batch scales to millions. Next: Hierarchical & DBSCAN — clustering without a preset K.

13 / 13