Hierarchical & Density-Based Clustering
K-Means needs K. K-Means needs spheres. Real data needs neither. Meet two powerful alternatives — algorithms that let the data decide its own shape and structure.
A Tree of Groupings
Hierarchical clustering builds a dendrogram — a full tree from n singletons to one mega-cluster. Cut at any height to read off any number of clusters. No K required before you start.
Merge Until One
- Start: every point is its own cluster (n clusters)
- Find the two closest clusters by linkage distance
- Merge them into one cluster
- Repeat until a single cluster remains
- Record each merge — this is the dendrogram
Four Ways to Measure Distance
- Single — min distance between any two points across clusters
- Complete — max distance; compact but can fragment large clusters
- Average — mean pairwise distance; robust compromise
- Ward — minimize variance increase; best in most cases
Ward’s Criterion
Ward’s method merges the pair that minimizes the increase in total within-cluster variance. Produces compact, balanced clusters. Default choice in most analyses.
Dendrogram Cuts
Merge height = dissimilarity at that step. A big jump in height between consecutive merges signals a natural boundary — cut there. The number of branches at the cut is your K.
Clusters Are Dense Regions
DBSCAN finds regions where points are tightly packed and separates them from sparse background. No K. Arbitrary shapes. Outliers labeled automatically as noise.
Core, Border, Noise
- Core — ≥ MinPts neighbors within ε; dense interior
- Border — within ε of a core point, but not dense itself
- Noise — not core, not near core; labeled as outlier
- Clusters = connected components of core points + their borders
ε-Neighborhood Definition
All points within distance ε form p’s neighborhood. If that neighborhood has ≥ MinPts members, p is a core point and seeds a cluster.
Arbitrary Shapes, Free Outliers
- No K — number of clusters emerges from data density
- Any shape — crescents, spirals, rings — no problem
- Outlier detection — noise points labeled automatically
- Deterministic — same input, same output every time
Choosing Your Algorithm
- K-Means — large datasets, spherical clusters, K known
- Hierarchical — interpretable tree, small/medium data, explore K
- DBSCAN — arbitrary shapes, outlier detection, K unknown
- When density varies greatly → consider HDBSCAN
What You Learned
Hierarchical builds a dendrogram; Ward’s linkage is best in practice. DBSCAN finds arbitrary-shaped clusters and labels outliers. Next: Dimensionality Reduction — PCA, t-SNE, UMAP.