Drive a polynomial fit from too-simple to too-complex in the sandbox, and predict — before you read it off the screen — the irreducible noise floor, how variance climbs with the polynomial degree, and how a ridge penalty pulls it back down.
The bias–variance tradeoff is the whole reason a more powerful model is not always a better one. A model too simple misses the signal (high bias²); a model too complex memorises the noise (high variance); and no model can ever beat the irreducible noise σ². Seeing those three terms add up to the test error is what makes overfitting concrete.
Open the sandbox. Set these four knobs and leave them there — every step below changes only the Degree and Ridge knobs. Using n = 60 points keeps the high-degree fits estimable, so the readouts never show “not estimable”.
Noise sigma -> 0.25 n (points) -> 60
Repeats -> 60 Seed -> 7 True function -> sine
Watch the bias², variance and noise σ² readouts under the panels — every number you predict is printed there to four decimals.
On the opening state, read the noise σ² readout. The noise term is exactly the variance of the measurement noise. With σ = 0.25, predict σ², then read it.
The noise readout is σ² = 0.25² = 0.0625. No model, however clever, can drive the test error below this floor — it is the irreducible term of the decomposition.
Set Degree d = 2 and n = 60, Ridge λ = 0. A quadratic cannot follow the sine, so bias² is large (about 0.1964) — but the fit barely moves when the data is resampled. Predict variance, then read it.
The variance readout is variance = 0.0243 — small, because a rigid quadratic (only three coefficients, edf = 3) changes little from one training set to the next. This is underfitting: bias dominates.
Raise Degree d to 8, still Ridge λ = 0. Now the fit chases every noisy point, so bias² collapses to about 0.0003. Predict what happened to variance versus degree 2, then read it.
The variance readout jumps to variance = 0.0370 — up from 0.0243 at degree 2. Bias fell and variance rose: that rising variance is overfitting, and you cannot see it in any single fitted curve.
Keep Degree d = 8 and turn the Ridge knob up to λ = 1. The penalty shrinks the large coefficients that made the fit thrash. Predict the new variance, then read it.
The variance readout drops to variance = 0.0091 — down from 0.0370 — while bias² rises only slightly (to about 0.0005). Regularization trades a little bias for far less variance: the λ knob is how you move along the tradeoff.
Everything above is waiting in the sandbox. Sweep the degree and watch the bias²-versus-variance readouts trade off, drag the ridge knob to move along the tradeoff, and turn on the decomposition panel to see where each term lives across the input.
- Read the irreducible noise floor, σ² = 0.0625
- Saw a too-simple fit hold variance to 0.0243 while bias dominated
- Watched variance climb to 0.0370 at degree 8 — overfitting made measurable
- Cut it back to 0.0091 with a ridge penalty λ = 1
- Every value you predicted is the demo's own bias–variance decomposition, not a picture