ML 101
M03 · L02
Strength in Numbers

Random Forests

One tree can overfit. Many trees working together — each trained on random data, each seeing random features — produce something remarkably powerful and robust.

01 / 14
ML 101
M03 · L02
The Problem

Trees Overfit

A deep decision tree memorizes training data. Change one data point and the whole structure changes. This is high variance — great on training data, terrible on new data.

100%
Train Acc
60%
Test Acc
02 / 14
ML 101
M03 · L02
The Insight

Wisdom of the Crowd

Individual trees make mistakes — but different trees make different mistakes. Average many trees and their errors cancel out. This is ensemble learning.

03 / 14
ML 101
M03 · L02
Step 1

Bootstrap Sampling

Each tree is trained on a different random sample of the training data, drawn with replacement. This is called bagging (Bootstrap Aggregating).

Key Property
On average, each bootstrap sample contains ~63% of unique training examples. The rest (~37%) are "out-of-bag" samples — used for validation.
04 / 14
ML 101
M03 · L02
Step 2

Random Feature Selection

At each split, each tree considers only a random subset of features — not all of them. This forces trees to be different from each other, breaking correlations.

Typical Feature Subsets
m = \begin{cases}\sqrt{p} & \text{classification}\\p/3 & \text{regression}\end{cases}
05 / 14
ML 101
M03 · L02
Step 3

Majority Vote

Each tree makes a prediction independently. The forest takes a majority vote (classification) or average (regression). More trees = more stable results.

🌲
Yes
🌲
Yes
🌲
No
🌲
Yes
→
YES
06 / 14
ML 101
M03 · L02
Free Validation

Out-of-Bag Error

Each sample is "out of bag" for ~37% of trees. Use those trees to predict it: no separate validation set needed. OOB error is a reliable estimate of generalization error.

37%
OOB Samples
Free
Validation
07 / 14
ML 101
M03 · L02
Bonus Insight

Feature Importance

Random forests measure how much each feature reduces impurity across all trees and all splits. Average this across the forest — you get a robust importance ranking of all features.

Mean Decrease in Impurity
\text{FI}(j) = \frac{1}{N_T}\sum_{t=1}^{N_T}\sum_{\text{splits on }j}\Delta\text{Gini}
08 / 14
ML 101
M03 · L02
Knobs to Turn

Key Hyperparameters

  • n_estimators — number of trees (more = better, diminishing returns)
  • max_features — features per split (√p for classification, p/3 for regression)
  • max_depth — depth of each tree (None = fully grown)
  • min_samples_split — min samples to split a node
  • bootstrap — whether to use bootstrap sampling (True)
09 / 14
ML 101
M03 · L02
Why It Works

Reducing Variance

Individual trees have low bias, high variance. Averaging N independent trees: variance shrinks by 1/N (if uncorrelated). Random feature selection ensures the trees are sufficiently uncorrelated.

10 / 14
ML 101
M03 · L02
Where It's Used

Random Forests in the Wild

  • Finance — fraud detection, credit scoring
  • Medicine — disease diagnosis, drug discovery
  • Remote sensing — land cover classification from satellite data
  • Ecology — species distribution modeling
  • Kaggle competitions — still a top baseline in 2024
11 / 14
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 14
ML 101
Key Takeaways
Summary

Key Takeaways

  • Random forests combine many trees via bagging + random features
  • Each tree sees a different bootstrap sample and feature subset
  • Final prediction = majority vote (classification) or mean (regression)
  • OOB error gives free validation without a separate set
  • Feature importance is a bonus — ranks predictors automatically
  • Reduces variance without increasing bias: the sweet spot
13 / 14
ML 101
Up Next
Coming Up

Gradient Boosting

Random forests build trees in parallel. Gradient boosting builds them sequentially — each new tree corrects the errors of all previous ones. Often even more accurate.

14 / 14