Take one problem through a whole ML workflow — first choose how complex the model should be and read the variance a ridge penalty leaves behind, then train it and read one weight after a single gradient-descent step, then evaluate it and read the precision off the confusion matrix. Predict each number before the screen shows it, and record your system at the end.
A model is never a single step. First you choose complexity: too simple underfits, too complex overfits, and a ridge penalty buys back variance at a little bias. Then you train: backpropagation computes each gradient and one step nudges every weight downhill, w ← w − η ∂L/∂w. Only then can you evaluate: a confusion matrix splits accuracy into precision and recall. This project is the interaction across those three stages — the workflow, not any one number.
Open each of the three sandboxes in a tab. Each step below names the one sandbox to read and the single choice to make in it; leave every other knob at its default so your numbers match these.
1 Bias-variance -> sigma = 0.25, n = 60, reps = 60, seed = 7, Degree d = 8, Ridge = 1
2 Backprop -> 2-2-1 sigmoid + MSE, eta = 1 3 Metrics -> prevalence 0.30, t = 0.50
Read three numbers across the three tabs: the variance readout in the bias-variance panel, the updated weight in the backprop grid, and the precision under the confusion matrix.
In the bias-variance sandbox pick a powerful model, Degree d = 8 on the sine target (σ = 0.25, n = 60). Left alone it overfits — variance is high. Now turn the Ridge penalty up to λ = 1 to rein it in. Predict the controlled variance, then read it.
The variance readout is variance = 0.0091 — a ridge penalty pulled a degree-8 fit's variance down from about 0.037 at little cost to bias. That is the complexity you carry forward to training. This is the demo's own bias-variance decomposition, not a typed number.
Switch to the backprop stepper and type the clean 2-2-1 sigmoid + MSE network: x = [1, 0], y = 1, W[1] = [[1,0],[0,0]], W[2] = [[1,1]]. The forward pass gives ŷ = 0.7740 and the backward pass ∂L/∂W[2]11 = -0.0578. With η = 1, predict the updated W[2]11 ← W[2]11 − η ∂L/∂w, then read it.
The update prints W[2]11 ← 1.0578 — the gradient was negative, so gradient descent nudged the weight up. One such step for every weight, thousands of times, is the whole of training. This is the demo's own forward and backward arithmetic, not a typed number.
Finally, in the metrics sandbox read how good the trained model is. On the opening state (prevalence 0.30, threshold t = 0.50) the confusion matrix reads TP = 273, FP = 74. Precision is TP / (TP + FP) — how often a positive prediction is right. Predict precision, then read it.
The precision readout is precision = 273 / (273 + 74) = 0.7867 — of every 100 sites the model flags, about 79 are genuinely positive. Accuracy alone would hide this; the confusion matrix is what makes evaluation honest. This is the demo's own confusion-matrix arithmetic, not a typed number.
Open all three sandboxes and run the workflow yourself: raise the degree and watch variance climb, then add ridge to pull it back, train the network a step at a time and watch the loss fall, then slide the threshold and watch precision and recall trade off. Each stage feeds the next.
An ML system is a set of choices and the numbers they produce at each stage. Write yours down — fill in each blank from the sandboxes, then add one sentence of reasoning. This is the deliverable; there is no number to read off the screen here.
Stage 1 Complexity ..... Degree d = 8, Ridge = 1 variance ...... ______
Stage 2 Training ....... 2-2-1 sigmoid, eta = 1 W[2]11 after 1 step .. ______
Stage 3 Evaluation ..... prevalence 0.30, t = 0.50 precision ..... ______
System note ............. does the evaluation justify the complexity you chose? _____
Design note ............. which stage would you change first, and why? __________
Then change one thing — a higher degree, a smaller learning rate, a rarer positive class — and note which of the three numbers moved and by how much. That coupling across choose, train and evaluate is the whole lesson of the project.