Module 3 · Lesson 2

Random Forests

12 min read
Article

A single decision tree is powerful but brittle. Train it with slightly different data and you get a completely different tree. This sensitivity — called high variance — is the core weakness of individual trees. Random forests solve this problem with a beautifully simple idea: build many trees, each slightly different, and let them vote.

Why One Tree Isn't Enough

A deep decision tree will memorize the training set almost perfectly. Given the same task with slightly perturbed data, it produces a very different structure. The model has low bias (it can fit complex patterns) but high variance (it's too sensitive to which specific data points it sees).

The numbers tell the story: a deep, unpruned tree might achieve 100% accuracy on training data and only 60% on test data. The gap between those numbers is the variance problem.

The key insight: If you have N independent models, each making random errors, the average of their predictions has variance N times smaller than any single model. Errors cancel out.

The Ensemble Idea: Wisdom of the Crowd

Suppose you ask 100 people to estimate the weight of a jar of coins. Each person is a bit off, but their errors go in different directions — some too high, some too low. Average their guesses and you'll likely be closer to the truth than any individual.

This is ensemble learning. Random forests apply the same principle to decision trees. Each tree is an imperfect but independent estimator. Their collective prediction is more reliable than any single tree's.

How Random Forests Work

Step 1: Bootstrap Sampling (Bagging)

Each tree in the forest is trained on a different dataset — but we don't collect new data. Instead, we sample from the existing training set with replacement. This is called a bootstrap sample.

If you have 1,000 training examples, each tree sees a random draw of 1,000 examples — but some appear twice, some three times, and about 37% don't appear at all. Different trees see different versions of the data, which forces them to learn differently.

This technique — training multiple models on bootstrap samples and combining them — is called bagging (Bootstrap Aggregating). It was introduced by Leo Breiman in 1994.

Step 2: Random Feature Selection

Bagging alone makes trees different in terms of which samples they see. But if all features are available at every split, trees trained on overlapping data will still look similar — especially near the root, where a single dominant feature tends to be chosen first.

Random forests add a second source of randomness: at each split in each tree, only a random subset of features is considered. The typical subset size follows a rule of thumb:

Feature Subset Size
m = \begin{cases}\sqrt{p} & \text{classification}\\p/3 & \text{regression}\end{cases}
where p is the total number of features. Classification problems use the square root; regression uses one-third.

By forcing each split to consider only a subset of features, the trees become decorrelated — they tend to use different features at different splits, making their errors more independent and more likely to cancel.

Step 3: Majority Vote (or Average)

Once all trees are trained, making a prediction is simple. Each tree independently classifies the input. For classification, the forest takes the majority vote — whichever class most trees predicted. For regression, it takes the mean across all trees.

Example: Four trees vote [Yes, Yes, No, Yes] → forest predicts Yes (3 out of 4). Adding more trees makes the vote more stable and reliable.

Out-of-Bag Error: Free Validation

Remember that ~37% of training examples are not used by any given tree. These are called out-of-bag (OOB) samples for that tree. We can use them as a validation set — evaluate each training example using only the trees that did not see it during training.

Aggregating these predictions gives the OOB error estimate — an unbiased estimate of generalization error, essentially for free, without needing a separate validation set or cross-validation. This is one of the practical beauties of random forests.

Feature Importance

A useful side effect of training random forests is a ranking of how important each feature is. The standard measure is the Mean Decrease in Impurity (MDI):

Feature Importance
\text{FI}(j) = \frac{1}{N_T}\sum_{t=1}^{N_T}\sum_{\text{splits on }j}\Delta\text{Gini}
FI(j) measures how much feature j reduces Gini impurity on average, summed over all splits that use feature j across all N_T trees.

Features that appear early in trees (near the root) and produce large purity gains receive high importance scores. The final scores are normalized so they sum to 1.

This gives you a free ranking of which variables matter most for prediction — useful for feature selection, model interpretation, and understanding your problem domain.

Key Hyperparameters

Random forests are remarkably easy to tune — they work well with default settings and are not sensitive to most hyperparameters.

n_estimators
Number of trees. More is always better but with diminishing returns. 100–500 is a good starting range.
max_features
Features per split. Default: √p (classification), p/3 (regression). Controls tree correlation.
max_depth
Maximum depth of each tree. Default is None (fully grown). Shallower trees reduce variance but increase bias.
min_samples_split
Minimum samples required to split an internal node. Higher values act as regularization.
bootstrap
Whether to use bootstrap sampling (True by default). Setting to False uses the full dataset for each tree.
oob_score
If True, computes and stores the OOB error estimate. Free validation without extra data.

Why It Works: The Bias-Variance View

The bias-variance tradeoff explains random forests cleanly. A deep decision tree has low bias — it can approximate almost any function — but high variance — it changes dramatically with training data.

When you average N independent models, the bias stays the same, but the variance drops by a factor of N. The catch is that trees are not fully independent — they're built from overlapping data and share features. The correlation between trees limits the variance reduction.

Random feature selection is the key innovation that reduces tree correlation. By forcing each tree to use different feature subsets, the trees become more independent, and the variance reduction approaches the theoretical 1/N limit.

Real-World Applications

Random forests are often the first model to try on structured/tabular data problems. They handle missing values, mixed feature types, and high-dimensional inputs gracefully, and rarely overfit badly.

Key Takeaways