Gradient Boosting
Random forests build trees in parallel, each independent. Gradient boosting builds them sequentially — each new tree corrects the errors left by all previous ones. A precise, powerful strategy.
Correct the Residual
Start with a simple prediction. Compute the residuals — what you got wrong. Train the next tree to predict those residuals. Add it to your model. Repeat until good enough.
Sequential Ensemble
The final model is the sum of all trees, each scaled by a learning rate η.
Gradient Descent in Function Space
Each tree fits the negative gradient of the loss function with respect to the current predictions. For MSE loss, this is exactly the residuals. For other losses, it generalizes elegantly.
The Learning Rate
Each tree is added with a small weight η (0.01–0.3 typically). This shrinkage prevents any single tree from dominating. Small η = more trees needed, but better generalization. Large η = fewer trees, risk of overfitting.
Shallow Trees
Boosting uses shallow trees — often depth 3–6. These are "weak learners" that only slightly improve over random guessing. Many weak learners combined sequentially become a very strong model.
XGBoost / LightGBM
XGBoost (2014) added regularization, parallelization, and smarter tree construction to classic gradient boosting. LightGBM (2017) added leaf-wise growth and histogram-based splits for 10–100× speedup. Together they dominate Kaggle competitions.
Regularized Objective
XGBoost minimizes a regularized objective that includes a complexity penalty on each tree, preventing overfitting by construction.
Tuning Gradient Boosting
- n_estimators — number of trees (use early stopping)
- learning_rate (η) — step size per tree (0.01–0.3)
- max_depth — tree depth (3–6 for most problems)
- subsample — fraction of samples per tree (0.6–0.9)
- colsample_bytree — fraction of features per tree
- reg_lambda / reg_alpha — L2 / L1 regularization
Early Stopping
Monitor validation loss at each boosting round. Stop adding trees when validation loss starts increasing. No need to grid-search n_estimators — just set it high and let early stopping find the right number automatically.
Boosting vs. Bagging
- RF builds in parallel, boosting builds sequentially
- RF reduces variance, boosting reduces bias
- RF is harder to overfit, boosting needs regularization
- Boosting usually wins on accuracy with proper tuning
- RF is faster to train, easier to parallelize
- Boosting dominates tabular data competitions
Key Takeaways
- Gradient boosting trains trees sequentially, each fitting residuals of the previous
- Each tree is scaled by learning rate η — smaller η = better generalization
- Uses shallow "weak learner" trees (depth 3–6)
- XGBoost / LightGBM add regularization + speed
- Use early stopping to find optimal n_estimators automatically
- Often the highest-accuracy method on tabular data
Module 4: Neural Networks
Trees handle tabular data brilliantly. But for images, text, and audio — the world moves to neural networks: layers of interconnected nodes that learn hierarchical representations directly from raw data.