A single decision tree is powerful but brittle. Train it with slightly different data and you get a completely different tree. This sensitivity — called high variance — is the core weakness of individual trees. Random forests solve this problem with a beautifully simple idea: build many trees, each slightly different, and let them vote.
A deep decision tree will memorize the training set almost perfectly. Given the same task with slightly perturbed data, it produces a very different structure. The model has low bias (it can fit complex patterns) but high variance (it's too sensitive to which specific data points it sees).
The numbers tell the story: a deep, unpruned tree might achieve 100% accuracy on training data and only 60% on test data. The gap between those numbers is the variance problem.
The key insight: If you have N independent models, each making random errors, the average of their predictions has variance N times smaller than any single model. Errors cancel out.
Suppose you ask 100 people to estimate the weight of a jar of coins. Each person is a bit off, but their errors go in different directions — some too high, some too low. Average their guesses and you'll likely be closer to the truth than any individual.
This is ensemble learning. Random forests apply the same principle to decision trees. Each tree is an imperfect but independent estimator. Their collective prediction is more reliable than any single tree's.
Each tree in the forest is trained on a different dataset — but we don't collect new data. Instead, we sample from the existing training set with replacement. This is called a bootstrap sample.
If you have 1,000 training examples, each tree sees a random draw of 1,000 examples — but some appear twice, some three times, and about 37% don't appear at all. Different trees see different versions of the data, which forces them to learn differently.
This technique — training multiple models on bootstrap samples and combining them — is called bagging (Bootstrap Aggregating). It was introduced by Leo Breiman in 1994.
Bagging alone makes trees different in terms of which samples they see. But if all features are available at every split, trees trained on overlapping data will still look similar — especially near the root, where a single dominant feature tends to be chosen first.
Random forests add a second source of randomness: at each split in each tree, only a random subset of features is considered. The typical subset size follows a rule of thumb:
By forcing each split to consider only a subset of features, the trees become decorrelated — they tend to use different features at different splits, making their errors more independent and more likely to cancel.
Once all trees are trained, making a prediction is simple. Each tree independently classifies the input. For classification, the forest takes the majority vote — whichever class most trees predicted. For regression, it takes the mean across all trees.
Example: Four trees vote [Yes, Yes, No, Yes] → forest predicts Yes (3 out of 4). Adding more trees makes the vote more stable and reliable.
Remember that ~37% of training examples are not used by any given tree. These are called out-of-bag (OOB) samples for that tree. We can use them as a validation set — evaluate each training example using only the trees that did not see it during training.
Aggregating these predictions gives the OOB error estimate — an unbiased estimate of generalization error, essentially for free, without needing a separate validation set or cross-validation. This is one of the practical beauties of random forests.
A useful side effect of training random forests is a ranking of how important each feature is. The standard measure is the Mean Decrease in Impurity (MDI):
Features that appear early in trees (near the root) and produce large purity gains receive high importance scores. The final scores are normalized so they sum to 1.
This gives you a free ranking of which variables matter most for prediction — useful for feature selection, model interpretation, and understanding your problem domain.
Random forests are remarkably easy to tune — they work well with default settings and are not sensitive to most hyperparameters.
The bias-variance tradeoff explains random forests cleanly. A deep decision tree has low bias — it can approximate almost any function — but high variance — it changes dramatically with training data.
When you average N independent models, the bias stays the same, but the variance drops by a factor of N. The catch is that trees are not fully independent — they're built from overlapping data and share features. The correlation between trees limits the variance reduction.
Random feature selection is the key innovation that reduces tree correlation. By forcing each tree to use different feature subsets, the trees become more independent, and the variance reduction approaches the theoretical 1/N limit.
Random forests are often the first model to try on structured/tabular data problems. They handle missing values, mixed feature types, and high-dimensional inputs gracefully, and rarely overfit badly.