AI-generated Computed, not drawn A Monte Carlo estimate, not an exact equality

This sandbox was generated by Claude (Anthropic) for the ML 101 course materials, and it does the thing most demonstrations of this topic only draw a cartoon of: it computes the bias–variance decomposition. Bias and variance are expectations over training sets — you cannot read either off a single fitted curve — so the page draws R independent training sets, fits every one by ridge least squares, and measures the average fitted function, the spread around it, and the noise floor. It then checks bias² + variance + σ² against a test error measured on fresh data, by a completely different route.

It is an estimate, and it says so. The residual in that identity is Monte Carlo error: it shrinks like 1/√R and never reaches zero. Measured in a stable setting it falls from about 3% at 200 repetitions to under 1% at 3200. And sometimes the expectation is not estimable at all — with few points, a high degree and no penalty, a small fraction of training sets produce catastrophic fits whose peak prediction is fifty times the median, and the mean is dominated by those rare draws. The page flags that state instead of printing an unconverged average as a result, and shows the cure: the ridge penalty the course teaches in m2-l3 brings the tail ratio from about 50 down to about 3.

Two numerical choices, stated because they could be mistaken for modelling ones. The basis is Chebyshev, not 1, x, x², …: it spans exactly the same space, so a degree-d Chebyshev fit is a degree-d polynomial fit with the same bias and variance, but the normal equations stay well conditioned where a Vandermonde matrix would not. And bias²/variance are integrated over x on a midpoint grid; an evenly spaced grid including x = ±1 gave the endpoints full weight, which overstated the identity's left side by 2.6× to 6× in exactly the unstable settings, always in the same direction.

Chosen rather than computed: the four true functions, the noise level, the x range, the default number of repetitions, the seed, and the tail-ratio cut-off of 8 above which the average is declared not estimable. The true function is deliberately never a polynomial — if it were, some finite degree would have exactly zero bias and the trade-off would be an artefact of the setup. Everything else on the page is computed live, including the effective degrees of freedom as trace(H) and the condition number of every system solved.

Bias–variance sandbox — what overfitting actually is

Course demo — linked from the Module 1 and Module 2 lesson decks in both languages; the page itself is English‑only for now. Built for ML-101 Lessons 1.3 and 2.3 and for the “adjust model complexity and see overfitting” half of the ML Playground lab. Both of those decks have zero canvases, so overfitting, the train/test split and “choosing λ” are described in words and never shown. The one thing to take away: variance is invisible in a single fit. It is how much the fitted curve moves when you resample the data, so this page fits many models and shows you the fan.

 

 

 

 

 
 

 

 
 

 

 
 

 

 
 

  

1

 

 

0

 

  

24

 

0.25

 

 

  

60

 

 

7

 

 

  

 

 — 
 — 
 — 
 — 
 — 
 — 
 — 
 — 

1 

 

 

2 

 

 

3 

 

 

4 

 

 

The arithmetic, in full

The experiment in words