ML 101
M03 · L01
Module 3

Decision Trees

What if your model reasoned like a human — asking a series of yes/no questions until it reached an answer? That is exactly what a decision tree does.

01 / 13
ML 101
M03 · L01
Anatomy

Tree Structure

A decision tree is a flowchart of binary decisions. Each internal node asks a question about a feature. Each branch is an answer. Each leaf is the final prediction.

Root Node
First split
Leaf Node
Prediction
02 / 13
ML 101
M03 · L01
Splitting Criterion

Gini Impurity

The Gini impurity measures how often a randomly chosen element would be incorrectly classified. A pure node (all one class) has Gini = 0. A perfectly mixed node has Gini = 0.5.

Gini Impurity
G = 1 - \sum_{k=1}^{K} p_k^2
03 / 13
ML 101
M03 · L01
Alternative Criterion

Entropy & Information Gain

Entropy measures disorder in a node. We choose the split that maximizes information gain — the reduction in entropy after the split. Both Gini and entropy produce similar trees in practice.

Entropy
H = -\sum_{k=1}^{K} p_k \log_2 p_k
04 / 13
ML 101
M03 · L01
Algorithm

Growing a Tree

  • 1. Start at the root with all training data
  • 2. Find the best feature + threshold to split
  • 3. Split data into two child nodes
  • 4. Repeat recursively on each child
  • 5. Stop when a stopping condition is met
05 / 13
ML 101
M03 · L01
The Problem

Overfitting Trees

A fully grown tree memorizes the training data, achieving zero training error but poor generalization. The tree has learned the noise, not the pattern.

Signs of Overfitting
Very deep tree • Many leaves with single samples • High training accuracy, low test accuracy
06 / 13
ML 101
M03 · L01
Regularization

Pruning the Tree

Pruning removes branches that add little predictive power. Pre-pruning stops growth early (max depth, min samples per leaf). Post-pruning grows fully then removes weak branches.

Pre-prune
Stop early
Post-prune
Trim after
07 / 13
ML 101
M03 · L01
Hyperparameters

Tuning Controls

  • max_depth — maximum tree levels
  • min_samples_split — minimum samples to split a node
  • min_samples_leaf — minimum samples in a leaf
  • max_features — features considered per split
  • criterion — gini or entropy
08 / 13
ML 101
M03 · L01
Strengths

Why Trees Shine

  • Interpretable: you can visualize and explain every decision
  • No scaling needed: features don’t need normalization
  • Mixed data: handles numeric and categorical features
  • Non-linear: captures complex decision boundaries
  • Fast inference: just traverse the tree
09 / 13
ML 101
M03 · L01
Weaknesses

The Limits

  • Unstable: small data changes can alter the tree completely
  • Axis-aligned: diagonal boundaries require many splits
  • Greedy: locally optimal splits aren’t globally optimal
  • Overfitting: easy to memorize training data
10 / 13
ML 101
M03 · L01
Visualization

A Simple Example

Classifying whether to play tennis based on weather. The root asks about Outlook, then branches ask about Humidity or Wind, until we reach a prediction.

Outlook? Humidity? Wind? Yes No Yes No Sunny Rain
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

Decision trees split data by asking feature questions, guided by Gini impurity or entropy. They are interpretable and flexible but prone to overfitting — cured by pruning. Next: Random Forests make them robust by averaging many trees.

Next Lesson
13 / 13