AI-generated Computed, not drawn Invariance holds in expectation, not in one sample

Generated by Claude (Anthropic) for the ML 101 course materials. This is deliberately not the “confusion-matrix threshold slider” the backlog asked for: m1-l4 slide 10 already is one, with a threshold slider, a ROC canvas and a live TP/FP/FN readout, and slide 5 already owns the confusion-matrix grid. Rebuilding it would duplicate a shipped widget, which this project's own rule forbids — if one slider makes the point, it belongs on a slide. So this page takes the dimension a threshold slider cannot reach, and the one the deck's “99% accurate, useless” slide gestures at without showing: prevalence and cost.

The headline is measured, not asserted. Holding the two class score distributions fixed and changing only the mix: AUC does not move; precision, F1 and PR-AUC collapse. Measured on this page's own generator, AUC runs 0.967 / 0.967 / 0.965 / 0.962 as prevalence falls 0.50→0.05, while PR-AUC goes 0.966→0.687 and precision 0.892→0.307. Any metric that moves in that experiment is telling you about your dataset, not your model. And a second identity is checked, not claimed: the deck says AUC is “the chance a random positive outscores a random negative”. Here the area under the ROC is computed by the trapezoid rule and by comparing every positive against every negative — two independent routes that agree to about 10−15.

The one published result: for a calibrated score the cost-optimal threshold is cFP/(cFP+cFN) — the standard decision-theoretic optimum, depending on the cost ratio and not on prevalence. It is printed beside a search of the real cost curve, and the scores here are a sigmoid of a Gaussian and therefore not calibrated, so the two can disagree. That gap is shown rather than glossed, and it is itself the argument for calibrating a model before taking decisions from its scores.

Chosen rather than computed: the three score shapes, the class separation and spread, the default prevalence and costs, the sample size and the seed. Stated rather than hidden: the prevalence invariance holds in expectation, not in any finite sample — at low prevalence there are few positives, so the AUC estimate itself turns noisy. The positive count is displayed on every row and flagged when it falls below 30.

Prevalence, cost and the threshold — which metrics measure your model

Course demo — linked from the Module 1 lesson deck in both languages; the page itself is English‑only for now. Built for ML-101 Lesson 1.4 and for the evaluation-metrics half of the ML Playground lab. The deck already lets you sweep a threshold; this page adds the two knobs it cannot reach. One thing to take away: ROC and AUC cannot see class prevalence, and precision, F1 and PR-AUC cannot ignore it. Hold the model fixed, change only how rare the positives are, and watch three of your four favourite numbers move.

 

 

 

 

 
 

 

 
 

 

 
 

 

 
 

  

0.30

 

 

2.4

 

0.9

 

 

  

0.50

 

 

  

1

 

5

 

 

  

1000

 

7

 

 — 
 — 
 — 
 — 
 — 
 — 
 — 
 — 

1 

 

 

2 

 

 

3 

 

 

4 

 

 

5 

 

 

6 

 

 

The arithmetic, in full

The chain in words