ML 101
M05 · L02
Module 5

Hierarchical & Density-Based Clustering

K-Means needs K. K-Means needs spheres. Real data needs neither. Meet two powerful alternatives — algorithms that let the data decide its own shape and structure.

01 / 13
ML 101
M05 · L02
Hierarchical

A Tree of Groupings

Hierarchical clustering builds a dendrogram — a full tree from n singletons to one mega-cluster. Cut at any height to read off any number of clusters. No K required before you start.

Bottom-up
Agglomerative
Top-down
Divisive
02 / 13
ML 101
M05 · L02
Algorithm

Merge Until One

  • Start: every point is its own cluster (n clusters)
  • Find the two closest clusters by linkage distance
  • Merge them into one cluster
  • Repeat until a single cluster remains
  • Record each merge — this is the dendrogram
03 / 13
ML 101
M05 · L02
Linkage

Four Ways to Measure Distance

  • Single — min distance between any two points across clusters
  • Complete — max distance; compact but can fragment large clusters
  • Average — mean pairwise distance; robust compromise
  • Ward — minimize variance increase; best in most cases
04 / 13
ML 101
M05 · L02
Best Linkage

Ward’s Criterion

Ward’s method merges the pair that minimizes the increase in total within-cluster variance. Produces compact, balanced clusters. Default choice in most analyses.

Ward Merge Cost
\Delta(A,B) = \dfrac{n_A n_B}{n_A+n_B}\,d(\boldsymbol{\mu}_A,\boldsymbol{\mu}_B)^2
05 / 13
ML 101
M05 · L02
Reading the Tree

Dendrogram Cuts

Merge height = dissimilarity at that step. A big jump in height between consecutive merges signals a natural boundary — cut there. The number of branches at the cut is your K.

Key Insight
Run once • Explore all K values • Choose after seeing the tree
06 / 13
ML 101
M05 · L02
Density-Based

Clusters Are Dense Regions

DBSCAN finds regions where points are tightly packed and separates them from sparse background. No K. Arbitrary shapes. Outliers labeled automatically as noise.

Param 1
ε radius
Param 2
MinPts
07 / 13
ML 101
M05 · L02
Point Types

Core, Border, Noise

  • Core — ≥ MinPts neighbors within ε; dense interior
  • Border — within ε of a core point, but not dense itself
  • Noise — not core, not near core; labeled as outlier
  • Clusters = connected components of core points + their borders
08 / 13
ML 101
M05 · L02
Formulation

ε-Neighborhood Definition

All points within distance ε form p’s neighborhood. If that neighborhood has ≥ MinPts members, p is a core point and seeds a cluster.

DBSCAN Neighborhood
N_{\varepsilon}(p) = \{\,q \in X \mid d(p,q) \leq \varepsilon\,\}
09 / 13
ML 101
M05 · L02
Why DBSCAN

Arbitrary Shapes, Free Outliers

  • No K — number of clusters emerges from data density
  • Any shape — crescents, spirals, rings — no problem
  • Outlier detection — noise points labeled automatically
  • Deterministic — same input, same output every time
10 / 13
ML 101
M05 · L02
When to Use

Choosing Your Algorithm

  • K-Means — large datasets, spherical clusters, K known
  • Hierarchical — interpretable tree, small/medium data, explore K
  • DBSCAN — arbitrary shapes, outlier detection, K unknown
  • When density varies greatly → consider HDBSCAN
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

Hierarchical builds a dendrogram; Ward’s linkage is best in practice. DBSCAN finds arbitrary-shaped clusters and labels outliers. Next: Dimensionality Reduction — PCA, t-SNE, UMAP.

13 / 13