Reading
Stories Mode

Evaluating Models

~18 min read Lesson 4 of Module 1

Is Your Model Any Good?

You can now describe what a model is, what it learns from, and how it learns. The previous lesson ended on the train/test split — the first honest thing you can do with a trained model. This lesson is the rest of that story: the numbers that tell you whether the model works, and the numbers that will fool you if you trust them alone.

Everything from here on assumes it. When a later lesson tells you to “tune C and gamma by grid search, scoring on F1”, or asks whether you want accuracy, AUC-ROC or RMSE, this is the lesson that makes that sentence mean something.

Core Idea

A metric is not a verdict, it is a question you chose to ask. Change the question and the same model looks brilliant or worthless. Picking the metric is a modelling decision, and it belongs to you, not to the library.

Same model, same data  →  accuracy 98.7%  |  precision 0.40  |  AUC 0.90

Three Sets, Not Two

A two-way train/test split is enough to detect overfitting. It is not enough to develop a model, because development means making choices: which algorithm, how many layers, how much regularisation, which features. Every one of those choices is fitted to whatever data you evaluate it on.

So real workflows use three disjoint sets. A common division is 60% train, 20% validation, 20% test, though the exact proportions matter far less than the separation:

Set Used for How often you look at it
Training set Fitting the parameters — the weights the optimiser adjusts Continuously, every iteration
Validation set Choosing between models and settings — the hyperparameters nobody optimises for you Often, once per candidate
Test set Estimating how the finished model will perform on data it has never influenced Once, at the very end

The validation set is the one most beginners skip, and it is the one that makes the other two honest. Its job is to absorb the damage of your search. You will fit dozens of models to the training data and rank them on validation; the winner's validation score is optimistic, because you chose it for that score. The test set exists precisely to give you an estimate that no choice of yours has contaminated.

Data Leakage: The Quiet Failure

Leakage is any information reaching the model that will not be available when it is deployed. It is quiet because it never produces an error — it produces a better-looking score, which is exactly what you were hoping for.

Three common forms. Preprocessing leakage: you compute the mean and standard deviation for standardisation over the whole dataset, then split. The training rows now carry a trace of the test rows. Duplicate leakage: the same customer, image or measurement appears in both train and test, so the model can recall rather than generalise. Temporal leakage: you shuffle time-series data randomly, and the model gets to see next week before predicting last week.

There is a fourth form that no code review catches, and it is the most common of all. You look at the test score, decide the model needs more regularisation, retrain, and look again. That loop is human gradient descent on the test set. Repeat it thirty times and your test score has become a validation score — still a number, no longer an estimate of anything.

The Discipline

Split first, then do everything else. Fit scalers, encoders and feature selectors on the training portion only, and apply them to the others. Tune on validation. Open the test set once, report the number, and if you do not like it, the honest move is to acknowledge that number — not to keep going.

Why Accuracy Lies

Accuracy is the fraction of predictions that were right. It is the metric everyone reaches for first, and on balanced problems it is perfectly reasonable.

Accuracy
\text{Accuracy}=\frac{TP+TN}{TP+TN+FP+FN}
TP, TN, FP and FN are the four counts defined in the next section. Accuracy weights all four equally, which is the source of the problem.

Now a concrete case. You are building a fault detector for a network of 10,000 cell sites. On any given day, 100 of them have a fault and 9,900 are healthy — 1% positive, 99% negative. This shape is not unusual; it is what fraud, disease screening, defect detection and equipment failure all look like.

Here is a model that requires no machine learning at all:

The 99% Model

def predict(site): return "healthy"

It is right 9,900 times out of 10,000. 99.0% accuracy. It finds zero faults. It would be beaten by a coin flip at the only thing you wanted it to do.

A genuinely useful detector on the same data scores 98.7% — worse than the useless one. If accuracy were your only metric you would ship the constant and delete the model. That is not a subtle statistical trap; it is a decision you would actually make, on evidence that looks conclusive.

The Confusion Matrix

Accuracy is one number squeezed out of four, and the two that matter get lost in the squeeze. The confusion matrix keeps all four. For a binary problem it is a 2×2 table of predicted class against actual class. Module 4 revisits it in the context of SVMs, where you will see confusion_matrix printed alongside a classification report — this section is where the four cells get their names.

Here is our real detector, at its default cut-off:

Cell Name Count What it costs
TP True positive — faulty, flagged 60 Nothing. This is the point.
FN False negative — faulty, missed 40 A broken site nobody is going to visit
FP False positive — healthy, flagged 90 An engineer drives out to a site that is fine
TN True negative — healthy, cleared 9,810 Nothing, and there are a great many of them

Accuracy is (60 + 9,810) / 10,000 = 98.7%. Notice what carries that number: TN, at 9,810, drowns out everything else. The 40 missed faults and the 90 wasted call-outs together move accuracy by just over one percentage point. The metric is dominated by the class you were not interested in.

So ignore TN for a moment and read the other three cells directly. That is exactly what the next three metrics do.

Precision: Can I Trust an Alarm?

Precision looks only at the predictions the model made positive, and asks how many of them were correct. It is the metric of the person receiving the alerts.

Precision
\text{Precision}=\frac{TP}{TP+FP}=\frac{60}{60+90}=0.40
Of the 150 sites the detector flagged, 60 were genuinely faulty. Precision 0.40 — three in five call-outs are wasted journeys.

Precision has no opinion about the faults you missed. A detector that flags exactly one site and gets it right has precision 1.00 while 99 faults sit unreported. That is not a flaw in the metric — it is the reason you never quote precision on its own.

Recall: How Many Did I Find?

Recall looks only at the cases that were actually positive, and asks how many the model caught. It is the metric of the person who owns the consequences of a miss. You will also see it called sensitivity, or the true positive rate — three names, one quantity.

Recall
\text{Recall}=\frac{TP}{TP+FN}=\frac{60}{60+40}=0.60
Of the 100 real faults, the detector found 60. Recall 0.60 — forty sites are broken and nobody has been told.

And recall has no opinion about false alarms. “Always predict faulty” achieves recall 1.00 with precision 0.01. Between them, precision and recall pin down both failure modes, which is why they travel as a pair.

F1: One Number for Both

Ranking twenty candidate models on two numbers is awkward, so F1 combines them. Crucially it is the harmonic mean, not the arithmetic one.

F1 Score
F_1=2\cdot\frac{\text{Precision}\cdot\text{Recall}}{\text{Precision}+\text{Recall}}
With precision 0.40 and recall 0.60, F1 = 2 × 0.24 / 1.00 = 0.48.

The choice of mean is doing real work. Take precision 1.00 and recall 0.02 — a detector that flags two sites and is right about both. The arithmetic mean is a respectable 0.51. The harmonic mean is 0.039. The harmonic mean cannot be rescued by one strong half; it sits close to the weaker of the two. That is precisely the behaviour you want from a summary of a trade-off.

Which One Should You Prefer?

There is no correct default, and any source that gives you one is skipping the interesting part. The question is always: which mistake is more expensive here?

Prefer precision when a false positive is costly. A spam filter that sends a real job offer to the junk folder has done more damage than one that lets a spam message through, because the user pays for the false positive and barely notices the false negative. Same for automated blocking, automated refunds, and anything that acts without a human in the loop.

Prefer recall when a false negative is costly. A screening test that misses a real case has failed at its one job; a false alarm costs a follow-up test. Same for safety interlocks, structural inspection, and our cell-site detector — a wasted engineer visit costs a morning, an unreported fault costs a day of outage.

Prefer F1 when both hurt and you need a single number to optimise. This is the usual situation during model selection, which is why f1_weighted shows up so often as a scoring argument. Just remember it is a convenience for ranking, not a statement that precision and recall matter equally in your problem — if one genuinely matters more, weight it, or optimise the one you care about subject to a floor on the other.

The Threshold You Never Chose

Here is something the four cells hide. Most classifiers do not output a class — they output a score between 0 and 1. Something then compares that score to a threshold to produce a label, and unless you said otherwise, that threshold is 0.5.

0.5 is a default, not a discovery. Nothing about your problem chose it. And every confusion matrix, every precision, every recall you have seen so far is a property of one threshold. Move it and all of them move.

Raise the threshold to 0.8 and our detector flags only its most confident 30 sites — precision climbs to about 1.00, recall falls to 0.30. Lower it to 0.2 and it flags over two thousand: recall rises to 0.85, precision collapses to 0.04. One model, one set of weights, wildly different behaviour. The interactive version of this lesson lets you drag that threshold and watch both numbers move together.

Building the ROC Curve

If a single threshold gives a single (precision, recall) pair, then sweeping the threshold across its whole range traces out a curve of the model's behaviour. Plot it with the right two axes and you get the ROC curve — receiver operating characteristic, a name inherited from radar operators in the 1940s and worth no further attention.

The Two Axes
\text{TPR}=\frac{TP}{TP+FN}\qquad \text{FPR}=\frac{FP}{FP+TN}
Vertical axis: true positive rate, which is recall. Horizontal axis: false positive rate, the fraction of healthy sites wrongly flagged. Neither denominator mixes the two classes, which is what makes the curve independent of how you weight them.

The construction is mechanical, and doing it by hand once is worth more than reading about it three times:

1. Score every example in the validation set. 2. Start with the threshold above every score: nothing is flagged, so TPR = 0 and FPR = 0. That is the bottom-left corner. 3. Lower the threshold one score at a time. Each step converts one example from predicted-negative to predicted-positive: if it was truly positive, TPR steps up; if truly negative, FPR steps right. 4. Continue until the threshold is below every score: everything is flagged, TPR = 1 and FPR = 1, the top-right corner. Join the points.

Read the geometry and it interprets itself. A model whose scores carry no information steps up and right in equal measure, tracing the diagonal. A model that ranks every positive above every negative goes straight up the left edge and along the top, through the corner at (0, 1) — perfect separation. Real models bow somewhere between. Up and to the left is better.

AUC: Scoring Every Threshold at Once

The curve is informative but it is a picture, and you cannot sort models by picture. AUC — the area under the ROC curve — collapses it to one number between 0 and 1. The diagonal has area 0.50; perfect separation has area 1.00. Our detector comes out at 0.90.

AUC has an interpretation that is much more useful than “area”, and it is worth memorising: AUC is the probability that a randomly chosen positive example is scored higher than a randomly chosen negative one. An AUC of 0.90 means that if you pick one faulty site and one healthy site at random, the detector rates the faulty one higher nine times in ten. It measures how well the model ranks, with the threshold taken out of the picture entirely.

That property makes AUC the natural metric when the threshold is not yet decided, or when it will be tuned later by whoever operates the system. It is also why AUC and F1 can disagree about which of two models is better: F1 judges one operating point, AUC judges the ordering.

AUC Flatters Imbalanced Problems

Our detector has AUC 0.90 and precision 0.40. Both are correct. FPR divides by 9,900 negatives, so 90 false positives read as an FPR of 0.009 — almost nothing. Precision divides by the 150 flagged sites, so the same 90 dominate. When the negative class is enormous, report AUC alongside precision and recall, never instead of them. The precision–recall curve, which swaps FPR for precision, is the sharper tool on data this skewed.

Regression Metrics: RMSE and MAE

None of the above applies when the model predicts a number rather than a class. There is no confusion matrix for “throughput will be 47.3 Mbps” — the prediction is not right or wrong, it is off by some amount, and the question becomes how you want to summarise those amounts.

You met MSE in the previous lesson as a loss function. As a reported metric it has an awkward property: its units are the square of what you predicted. “MSE of 10.9 squared-megabits-per-second” is not a sentence anyone can act on. Taking the square root fixes that, giving RMSE.

Root Mean Squared Error
\text{RMSE}=\sqrt{\frac{1}{n}\sum_{i=1}^{n}(\hat{y}_i-y_i)^2}
Back in the original units. RMSE is always at least as large as MAE, and larger the more unevenly the errors are spread.

MAE takes the average of the absolute errors instead of the squared ones. It answers the plainest possible question — how far off is this model, typically?

Mean Absolute Error
\text{MAE}=\frac{1}{n}\sum_{i=1}^{n}\left|\hat{y}_i-y_i\right|
Every error contributes in proportion to its size — no more, no less.

The difference between them is entirely about outliers, and one worked example makes it concrete. Predict throughput at ten sites. Nine predictions are off by 1 Mbps; the tenth is off by 10 Mbps.

MAE = (9 × 1 + 10) / 10 = 1.9 Mbps. RMSE = √((9 × 1 + 100) / 10) = √10.9 = 3.3 Mbps. The single bad prediction contributes 100 of the 109 squared error — 92% of RMSE comes from 10% of the data. MAE gives that same prediction one tenth of the weight.

R²: Compared to What?

RMSE and MAE are absolute. An RMSE of 3.3 Mbps is good if throughput varies over hundreds and terrible if it varies over five, and the number alone cannot tell you which. R² — the coefficient of determination — supplies the missing comparison by measuring you against the laziest possible model: always predict the mean.

Coefficient of Determination
R^2=1-\frac{\sum_{i}(\hat{y}_i-y_i)^2}{\sum_{i}(\bar{y}-y_i)^2}
The numerator is your model's squared error; the denominator is the squared error of predicting the mean of y every time. R² is the fraction of that baseline error you removed.

Our ten sites had a squared error of 109 against 1,000 for always predicting the mean, so R² = 1 − 109/1,000 = 0.89. Read it as: the model explains 89% of the variation the mean leaves unexplained. R² = 1 is perfect, R² = 0 means you have exactly matched the mean, and R² can go negative — a model worse than the mean, which happens more often than you would expect on a test set.

When MAE Is the Honest Choice

RMSE is the default in most papers and most tutorials, so it is worth being clear about when it is the wrong report. Ask one question: is a large error genuinely worse, in proportion to its square?

Sometimes it is. If you are sizing a battery, a link budget or a buffer, being off by 10 really is much worse than ten errors of 1, because the single large miss is the one that breaks the system. Squaring encodes that, and RMSE is the right metric.

Often it is not. If your extreme errors come from corrupt data rather than bad predictions — a stuck sensor, a mislabelled row, a measurement taken during maintenance — then RMSE is reporting the quality of your data collection, not the quality of your model. Optimise it and the model contorts itself to chase noise it should be ignoring.

The practical rule: report both. They cost nothing to compute, and the gap between them is itself information — a RMSE much larger than MAE says your errors are dominated by a few large ones, and you should go and look at them.

The Evaluation Checklist

1. Split into train / validation / test before touching anything else. 2. Check the class balance — if it is skewed, accuracy is off the table. 3. Print the confusion matrix, not just a score. 4. Decide which mistake costs more, and pick your metric from that. 5. Check the threshold; 0.5 is a default, not an answer. 6. For regression, report RMSE and MAE together, plus R² for context. 7. Open the test set once.

Module 1 ends here. Module 2 returns to linear regression in earnest — and now that you can measure a model properly, you will be able to tell whether the improvements it introduces are real.

Key Takeaways
  • Three sets, not two: train fits the parameters, validation chooses between models, test is opened once and estimates real performance.
  • Leakage is any information reaching the model that will not exist at deployment — including your own repeated glances at the test score.
  • Accuracy is dominated by the majority class. On a 99%-negative problem, “always predict negative” scores 99.0% and finds nothing.
  • The confusion matrix keeps all four counts. Precision = TP/(TP+FP) is the trustworthiness of an alarm; recall = TP/(TP+FN) is the fraction of real cases found.
  • Prefer precision when false alarms are costly, recall when misses are costly, F1 — the harmonic mean — when you need one number to rank by.
  • The ROC curve is built by sweeping the threshold and plotting TPR against FPR. AUC is the probability a random positive outscores a random negative; 0.50 is chance.
  • RMSE punishes large errors quadratically, MAE weights every error in proportion to its size, and R² compares you to predicting the mean. Report RMSE and MAE together.
PreviousYour First ML Model Module Overview Next LessonLinear Regression