ML 101
M1 · Lab
Hands-on lab
Watch precision and recall pull apart

Slide the decision threshold, then change the class prevalence in the sandbox, and predict — before you read it off the confusion matrix — the precision, recall and F1 at a stated threshold, how raising the threshold trades recall away for precision, and how a rarer positive class quietly collapses precision while the AUC barely moves.

Two numbers from the confusion matrix
\text{precision} = \dfrac{TP}{TP + FP} \qquad \text{recall} = \dfrac{TP}{TP + FN}

A single accuracy number hides the two mistakes that matter. The confusion matrix splits them apart: precision asks how often an alarm is right, recall asks how many real positives you caught, and F1 is their harmonic mean. The threshold slides you along the tradeoff between them — and because precision counts predicted positives, it depends on how common the positive class is, which is why a model that looks excellent on a balanced test can be near-useless on a rare one.

1 / 8
ML 101
M1 · Lab
Set it up
One threshold

Open the sandbox. Set these knobs and leave them there — the first three steps change only the Threshold t slider, and the last step reads the built-in prevalence comparison. The defaults below are the boot state, so you can just confirm them.

Set these values
Prevalence -> 0.30 Threshold t -> 0.50 Separation -> 2.4 Spread -> 0.9 Shape -> normal n -> 1000 Seed -> 7

Watch the confusion matrix and the metric readouts under it — every number you predict (precision, recall, F1) is printed there, and the prevalence table on the same page shows what happens as the positive class gets rarer.

2 / 8
ML 101
M1 · Lab
Step 1 of 4
The confusion matrix, at t = 0.50

On the opening state (prevalence 0.30, threshold t = 0.50), read the confusion matrix: TP = 273, FP = 74, FN = 27, TN = 626. Precision counts how often a positive prediction is right, TP / (TP + FP). Predict precision, then read it.

Expected

The precision readout is precision = 273 / (273 + 74) = 0.7867. Of every 100 sites the detector flags, about 79 are genuinely positive — a number about the predicted positives, not about the whole test set.

3 / 8
ML 101
M1 · Lab
Step 2 of 4
Raise the threshold: recall drops

Drag Threshold t up to 0.80. Now the detector only fires when it is very sure, so almost every alarm is right (precision climbs to about 0.985) — but it stays silent on the harder true positives. Recall is TP / (TP + FN). Predict recall at t = 0.80, then read it.

Expected

The recall readout falls to recall = 0.4367 — down from 0.91 at t = 0.50. You now miss well over half the real positives. That is the precision-recall tradeoff: the threshold buys one only by spending the other.

4 / 8
ML 101
M1 · Lab
Step 3 of 4
F1: one number for both

Drag Threshold t back to 0.50. F1 is the harmonic mean of precision and recall, F1 = 2·P·R / (P + R), so it sits below their arithmetic average and only rewards a model that is good at both. With precision = 0.7867 and recall = 0.91, predict F1, then read it.

Expected

The F1 readout is F1 = 0.8439 — between precision and recall, and pulled toward the smaller of the two. One balanced score, but only honest when a false positive and a false negative cost about the same.

5 / 8
ML 101
M1 · Lab
Step 4 of 4
The base-rate trap

Now read the prevalence comparison table (the “what prevalence does to each metric” panel). It holds the model fixed and only makes the positive class rarer, walking prevalence down to 0.02. At prevalence 0.50 precision is about 0.89; predict what precision becomes at prevalence 0.02 (same threshold t = 0.50), then read it.

Expected

Precision collapses to 0.1382 — while AUC barely moves (it holds near 0.94 to 0.97 across the whole ladder). The model did not get worse; the same false-positive rate now lands on a far larger negative class, so most alarms are false. This is why precision must always be read against the base rate.

6 / 8
ML 101
M1 · Lab
Your turn
Open the sandbox

Everything above is waiting in the sandbox. Sweep the threshold and watch precision and recall trade off in real time, change the shape of the score distributions, and walk the prevalence ladder to see AUC hold steady while precision and PR-AUC fall away beneath it.

7 / 8
ML 101
M1 · Lab
Wrap-up
What you did
  • Read precision = 0.7867 off the confusion matrix at t = 0.50
  • Watched raising the threshold to 0.80 drop recall to 0.4367 — the precision-recall tradeoff
  • Combined both into F1 = 0.8439, the harmonic mean
  • Saw precision collapse to 0.1382 at prevalence 0.02 while AUC held — the base-rate trap
  • Every value you predicted is the demo's own confusion-matrix arithmetic, not a picture
8 / 8