Evaluating Models
You can train a model. Now the harder question: is it any good? This lesson gives you the numbers that answer that — and shows you why the most obvious one lies.
Three sets, not two
Last lesson you split data in two. Real work needs three, because you make two different kinds of decision and each needs its own clean data.
Never tune on test
Every time you look at the test score and change something, a little of that set leaks into your model. Do it thirty times and the test set has quietly become a second training set.
99% accurate, useless
10,000 cell sites. 100 have a fault. Write one line — “always predict healthy” — and you score 99.0% accuracy while finding not a single fault.
The confusion matrix
Accuracy collapses four numbers into one and hides the two that matter. Here is the real detector at its default threshold:
Precision
When the model raises an alarm, how often is it right? Our detector flagged 150 sites and 60 were genuinely faulty — precision 0.40. Six of every ten call-outs are wasted.
Recall
Of the faults that were actually there, how many did we find? 60 of 100 — recall 0.60. Forty broken sites are still broken and nobody has been told.
The F1 score
One number for both. It is the harmonic mean, not the plain average — so it refuses to be flattered by one strong half. Ours: 0.48.
Which one do you want?
There is no default. Ask which mistake costs more:
- Precision — when a false alarm is expensive. Spam filters: losing a real email is worse than one spam getting through.
- Recall — when a miss is expensive. Screening: a missed case is worse than a second test.
- F1 — when both hurt and you need one number to rank models by.
Sweep the threshold
The model outputs a score, not a label. You pick the cut-off. Drag it and watch the point trace the ROC curve.
Area under the curve
AUC is the shaded area you just watched. It has a plain meaning: the chance that a random faulty site scores higher than a random healthy one. Ours is 0.90; a coin-flip is 0.50.
RMSE or MAE?
Predicting a number, not a class. Both report error in the units you predicted in — but RMSE squares first, so one huge miss dominates it. MAE does not flinch.
MAE and R²
Nine predictions off by 1 and one off by 10: MAE is 1.9, RMSE is 3.3. If those outliers are bad data rather than bad predictions, MAE is the honest report.
What you learned
Train fits, validation chooses, test reports — once. Accuracy hides imbalance; the confusion matrix exposes it. Precision, recall and F1 name the trade-off, ROC and AUC score it across every threshold. For regression: RMSE punishes outliers, MAE tolerates them, R² compares you to guessing the mean.