Clustering — K-Means
No labels. No teacher. Just data and structure. Unsupervised learning finds hidden groupings — and K-Means is the algorithm that launched a thousand segmentations.
Grouping Without Labels
Clustering partitions data into groups so that within-group similarity is high and between-group similarity is low. No ground truth required — the structure emerges from the data itself.
Assign → Update → Repeat
- Initialize — place K centroids (randomly or via K-Means++)
- Assign — each point joins its nearest centroid
- Update — recompute each centroid as the cluster mean
- Check — stop if assignments did not change
- Guaranteed to converge; converges to a local minimum
Minimizing Inertia
K-Means minimizes the total within-cluster sum of squared distances from each point to its centroid. Lower inertia = tighter, more compact clusters.
Choosing K
- Elbow method — plot inertia vs. K; look for the bend
- Silhouette score — measures cohesion vs. separation per point
- Range −1 (wrong cluster) to +1 (perfect); maximize over K
- Domain knowledge often most reliable guide
- Neither metric is definitive — treat as heuristics
Silhouette Score
a(i) = avg distance to same-cluster points. b(i) = avg distance to nearest other cluster. s(i) near +1 means well-clustered; near 0 means borderline.
K-Means++
Random initialization risks poor local minima. K-Means++ seeds centroids by choosing each new one with probability proportional to its squared distance from the nearest existing centroid — spreading them out intelligently.
When K-Means Fails
- Spherical assumption — can't find curved or elongated clusters
- Equal sizes — large clusters get split; small ones merged
- Outlier sensitivity — one extreme point pulls a centroid far away
- K required upfront — must guess or search
- Euclidean only — not suited for categorical or text data
Mini-Batch K-Means
Standard K-Means loads all data per iteration. Mini-Batch K-Means updates centroids on a random subsample each step — like SGD for clustering. Dramatically faster on millions of points.
Scale Before Clustering
K-Means uses Euclidean distance. A feature with large values dominates all distance calculations. Always standardize to zero mean, unit variance before running K-Means.
Where K-Means Shines
- Customer segmentation — group users by behavior
- Image compression — reduce colors to K representative values
- Document clustering — topic discovery in large corpora
- Anomaly detection — points far from any centroid are outliers
- Feature engineering — cluster membership as a new feature
What You Learned
K-Means iterates assign→update to minimize inertia. K-Means++ seeds smartly. The elbow and silhouette help choose K. Mini-Batch scales to millions. Next: Hierarchical & DBSCAN — clustering without a preset K.