ML 101
M01 · L04
Introduction

Evaluating Models

You can train a model. Now the harder question: is it any good? This lesson gives you the numbers that answer that — and shows you why the most obvious one lies.

01 / 15
ML 101
M01 · L04
The setup

Three sets, not two

Last lesson you split data in two. Real work needs three, because you make two different kinds of decision and each needs its own clean data.

Train 60%
Fit
Validate 20%
Choose
Test 20%
Report
02 / 15
ML 101
M01 · L04
The rule

Never tune on test

Every time you look at the test score and change something, a little of that set leaks into your model. Do it thirty times and the test set has quietly become a second training set.

Leakage, plainly
Any information from outside the training set that reaches the model. The test set is opened once, at the end, and then it is spent.
03 / 15
ML 101
M01 · L04
The trap

99% accurate, useless

10,000 cell sites. 100 have a fault. Write one line — “always predict healthy” — and you score 99.0% accuracy while finding not a single fault.

Always healthy
99.0%
Real detector
98.7%
04 / 15
ML 101
M01 · L04
Four numbers

The confusion matrix

Accuracy collapses four numbers into one and hides the two that matter. Here is the real detector at its default threshold:

TP — caught
60
FN — missed
40
FP — false alarm
90
TN — correctly clear
9,810
05 / 15
ML 101
M01 · L04
Metric 1

Precision

When the model raises an alarm, how often is it right? Our detector flagged 150 sites and 60 were genuinely faulty — precision 0.40. Six of every ten call-outs are wasted.

Precision
\text{Precision} = \frac{TP}{TP + FP} = \frac{60}{60 + 90} = 0.40
06 / 15
ML 101
M01 · L04
Metric 2

Recall

Of the faults that were actually there, how many did we find? 60 of 100 — recall 0.60. Forty broken sites are still broken and nobody has been told.

Recall
\text{Recall} = \frac{TP}{TP + FN} = \frac{60}{60 + 40} = 0.60
07 / 15
ML 101
M01 · L04
Metric 3

The F1 score

One number for both. It is the harmonic mean, not the plain average — so it refuses to be flattered by one strong half. Ours: 0.48.

F1 score
F_1 = 2 \cdot \frac{P \cdot R}{P + R} = 2 \cdot \frac{0.24}{1.00} = 0.48
08 / 15
ML 101
M01 · L04
The choice

Which one do you want?

There is no default. Ask which mistake costs more:

  • Precision — when a false alarm is expensive. Spam filters: losing a real email is worse than one spam getting through.
  • Recall — when a miss is expensive. Screening: a missed case is worse than a second test.
  • F1 — when both hurt and you need one number to rank models by.
09 / 15
ML 101
Interactive
Try it

Sweep the threshold

The model outputs a score, not a label. You pick the cut-off. Drag it and watch the point trace the ROC curve.

Cut-off 0.50
TP 60 · FP 90 · FN 40
precision 0.40 · recall 0.60
10 / 15
ML 101
M01 · L04
One number, all thresholds

Area under the curve

AUC is the shaded area you just watched. It has a plain meaning: the chance that a random faulty site scores higher than a random healthy one. Ours is 0.90; a coin-flip is 0.50.

The two axes
\text{TPR} = \frac{TP}{TP + FN} \qquad \text{FPR} = \frac{FP}{FP + TN}
11 / 15
ML 101
M01 · L04
Regression

RMSE or MAE?

Predicting a number, not a class. Both report error in the units you predicted in — but RMSE squares first, so one huge miss dominates it. MAE does not flinch.

Root mean squared error
\text{RMSE} = \sqrt{\frac{1}{n}\sum_{i=1}^{n}(\hat{y}_i - y_i)^2}
12 / 15
ML 101
M01 · L04
Regression

MAE and R²

Nine predictions off by 1 and one off by 10: MAE is 1.9, RMSE is 3.3. If those outliers are bad data rather than bad predictions, MAE is the honest report.

Mean absolute error
\text{MAE} = \frac{1}{n}\sum_{i=1}^{n}\left|\hat{y}_i - y_i\right|
13 / 15
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson, using its own cell-site numbers. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

14 / 15
ML 101
Summary
Recap

What you learned

Train fits, validation chooses, test reports — once. Accuracy hides imbalance; the confusion matrix exposes it. Precision, recall and F1 name the trade-off, ROC and AUC score it across every threshold. For regression: RMSE punishes outliers, MAE tolerates them, R² compares you to guessing the mean.

Next Lesson
15 / 15