AI-generated Computed, not drawn Greedy: locally optimal, not globally

This sandbox was generated by Claude (Anthropic) for the ML 101 course materials. It builds a real decision tree on a small 2D dataset and shows the three things the Module 3 lesson describes in words: the Gini impurity and entropy of every node, the information gain of every candidate split, and the piecewise-constant decision boundary the finished tree induces. Nothing here is sketched — the boundary is drawn by running the tree's own predict() over a grid of points.

The split search is greedy, and the page shows what that costs. At each node the tree scores every axis-aligned split by information gain and keeps the single best one, never looking ahead. On the XOR dataset a depth‑1 tree therefore finds no useful split at all — every straight cut leaves both halves mixed — so its boundary is one flat colour and its accuracy is a coin flip. Allow depth 2 and it snaps into place. That is not a defect; it is the definition of greedy, made visible.

Gini and entropy usually agree. They are different measures of node impurity, but on most splits they rank the candidates the same way and pick the same cut — the page computes the best split under each criterion and tells you whether they landed on the same one. Where they differ it is by a hair, and the difference almost never changes the tree.

Computed: every dataset (seeded and reproducible), the class counts and impurity at every node, the information gain of every candidate split, the greedy tree, its node and leaf counts and depth, the train accuracy, and the decision boundary by evaluating predict() on a grid. Chosen rather than computed: the four dataset shapes, the label-noise level, the point count, the seed, and the default depth, min-samples and criterion. Raise the depth on a noisy dataset and watch train accuracy climb toward 100% as the tree memorises the noise — the overfitting the course warns about, computed on screen.

Decision-tree sandbox — impurity, gain, and the boundary that follows

Course demo — linked from the Module 3 lesson deck; the page itself is English‑only for now. A decision tree makes decisions by asking one yes/no question about one feature at a time. This page lets you set the depth, the stopping rule and the impurity criterion, and then computes the whole tree and its decision boundary from the data. One thing to take away: the boundary is axis-aligned rectangles by construction — a tree can only cut straight across one feature at a time, which is exactly why it needs depth to describe a diagonal or an XOR.

 

 

 

 

 
 

 

 
 

 

 
 

 

 
 

  

4

 

 

2

 

 

 

  

 

0.05

 

120

 

7

 

  

 

 

 — 
 — 
 — 
 — 
 — 
 — 
 — 

1 

 

 

2 

 

 

3 

 

 

4 

 

 

The arithmetic, in full

The tree, step by step