ML 101
M03 · L03
Learn from Mistakes

Gradient Boosting

Random forests build trees in parallel, each independent. Gradient boosting builds them sequentially — each new tree corrects the errors left by all previous ones. A precise, powerful strategy.

01 / 14
ML 101
M03 · L03
The Concept

Correct the Residual

Start with a simple prediction. Compute the residuals — what you got wrong. Train the next tree to predict those residuals. Add it to your model. Repeat until good enough.

02 / 14
ML 101
M03 · L03
Building Up

Sequential Ensemble

The final model is the sum of all trees, each scaled by a learning rate η.

T₁: base prediction
→
fit residuals
T₂: correct T₁ errors
→
fit residuals
T₃: correct T₁+T₂ errors
→
fit residuals
⋮
03 / 14
ML 101
M03 · L03
The Math

Gradient Descent in Function Space

Each tree fits the negative gradient of the loss function with respect to the current predictions. For MSE loss, this is exactly the residuals. For other losses, it generalizes elegantly.

Update Rule
F_m(x)=F_{m-1}(x)+\eta\,h_m(x)
04 / 14
ML 101
M03 · L03
Shrinkage

The Learning Rate

Each tree is added with a small weight η (0.01–0.3 typically). This shrinkage prevents any single tree from dominating. Small η = more trees needed, but better generalization. Large η = fewer trees, risk of overfitting.

0.1
Typical η
100+
Trees needed
05 / 14
ML 101
M03 · L03
Weak Learners

Shallow Trees

Boosting uses shallow trees — often depth 3–6. These are "weak learners" that only slightly improve over random guessing. Many weak learners combined sequentially become a very strong model.

Why Shallow?
Deep trees would overfit the residuals. Shallow trees capture only the most important signal at each step, leaving room for future trees to help.
06 / 14
ML 101
M03 · L03
Modern Variant

XGBoost / LightGBM

XGBoost (2014) added regularization, parallelization, and smarter tree construction to classic gradient boosting. LightGBM (2017) added leaf-wise growth and histogram-based splits for 10–100× speedup. Together they dominate Kaggle competitions.

07 / 14
ML 101
M03 · L03
XGBoost Objective

Regularized Objective

XGBoost minimizes a regularized objective that includes a complexity penalty on each tree, preventing overfitting by construction.

XGBoost Objective
\mathcal{L}=\sum_i l(y_i,\hat{y}_i)+\sum_k\Omega(f_k)
08 / 14
ML 101
M03 · L03
Key Parameters

Tuning Gradient Boosting

  • n_estimators — number of trees (use early stopping)
  • learning_rate (η) — step size per tree (0.01–0.3)
  • max_depth — tree depth (3–6 for most problems)
  • subsample — fraction of samples per tree (0.6–0.9)
  • colsample_bytree — fraction of features per tree
  • reg_lambda / reg_alpha — L2 / L1 regularization
09 / 14
ML 101
M03 · L03
Avoiding Overfitting

Early Stopping

Monitor validation loss at each boosting round. Stop adding trees when validation loss starts increasing. No need to grid-search n_estimators — just set it high and let early stopping find the right number automatically.

Typical Recipe
n_estimators=5000, learning_rate=0.05, early_stopping_rounds=50. Train on 80%, validate on 20%. Best round is often 200–800 trees.
10 / 14
ML 101
M03 · L03
vs. Random Forest

Boosting vs. Bagging

  • RF builds in parallel, boosting builds sequentially
  • RF reduces variance, boosting reduces bias
  • RF is harder to overfit, boosting needs regularization
  • Boosting usually wins on accuracy with proper tuning
  • RF is faster to train, easier to parallelize
  • Boosting dominates tabular data competitions
11 / 14
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 14
ML 101
Key Takeaways
Summary

Key Takeaways

  • Gradient boosting trains trees sequentially, each fitting residuals of the previous
  • Each tree is scaled by learning rate η — smaller η = better generalization
  • Uses shallow "weak learner" trees (depth 3–6)
  • XGBoost / LightGBM add regularization + speed
  • Use early stopping to find optimal n_estimators automatically
  • Often the highest-accuracy method on tabular data
13 / 14
ML 101
Up Next
Coming Up

Module 4: Neural Networks

Trees handle tabular data brilliantly. But for images, text, and audio — the world moves to neural networks: layers of interconnected nodes that learn hierarchical representations directly from raw data.

14 / 14