ML 101
M10 · L02
Practical ML

Model Selection & Tuning

Choosing the right model and finding its best hyperparameters — the systematic science behind squeezing real-world performance out of your ML pipeline.

01 / 13
ML 101
M10 · L02
No Free Lunch

No Single Winner

Averaged over all possible problems, every algorithm performs equally. There is no universally best model — the right choice depends on your data, size, features, and constraints.

Practical Implication
Always compare 2–3 algorithm families on your specific problem before diving into hyperparameter tuning of any one model.
02 / 13
ML 101
M10 · L02
Honest Evaluation

Cross-Validation

A single train/test split can be lucky or unlucky. K-fold cross-validation rotates the held-out set across K folds and averages the scores:

K-Fold CV Score
\text{CV Score}=\frac{1}{K}\sum_{k=1}^{K}\text{Score}_k
03 / 13
ML 101
M10 · L02
CV Strategies

Match CV to Your Data Type

  • Stratified K-Fold — preserves class ratios. Use for classification
  • Time Series Split — expanding window: always train on past, validate on future
  • Group K-Fold — keep all rows from one entity (user, patient) in the same fold
  • Leave-One-Out — K = N; for very small datasets only
04 / 13
ML 101
M10 · L02
Hyperparameter Search

Grid Search vs. Random Search

  • Grid search — tries every combination. Guaranteed best on grid, but exponential cost
  • Random search — samples configurations randomly. Covers more values of important parameters for the same budget
  • For ≤3 parameters with small grids: grid search. Otherwise: random search
  • Bergstra & Bengio (2012): random search finds better results than grid search in practice
05 / 13
ML 101
M10 · L02
Why Random Wins

Coverage vs. Exhaustion

With n trials and d parameters, grid search projects only n^(1/d) distinct values per axis. Random search projects n distinct values per axis — far more when d is large.

Intuition
Most hyperparameter spaces have low effective dimensionality — only a few parameters matter. Random search explores more of those critical dimensions.
06 / 13
ML 101
M10 · L02
Informed Search

Bayesian Optimization

Builds a surrogate model of the objective and picks the next configuration by balancing exploitation (near current best) and exploration (uncertain regions).

Use when evaluations are expensive (deep learning, long training). Tools: Optuna, Hyperopt, scikit-optimize.

07 / 13
ML 101
M10 · L02
Ensemble Methods

Combine for Strength

  • Voting / Averaging — train diverse models independently, average predictions
  • Stacking — meta-learner trained on out-of-fold predictions of base models
  • Blending — meta-learner trained on a fixed holdout set. Simpler but wastes data
  • Key insight: diversity beats accuracy for individual base models
08 / 13
ML 101
M10 · L02
Stacking Detail

Out-of-Fold Predictions

Never feed the meta-learner predictions from base models trained on the same rows. Use out-of-fold (OOF) predictions: for each fold, train on others, predict on this fold.

OOF
Leak-free predictions
Meta
Learns combination
+1–3%
Typical gain
09 / 13
ML 101
M10 · L02
AutoML

Automated Pipelines

  • Auto-sklearn — Bayesian opt over scikit-learn pipelines
  • TPOT — genetic programming; exports clean code
  • H2O AutoML — stacks GBMs and neural nets; production-ready
  • AutoGluon — strong tabular defaults from AWS
  • Use as a baseline and benchmark — understand what it does
10 / 13
ML 101
M10 · L02
The Full Workflow

A Systematic Process

  • Define your metric — what exactly are you optimizing?
  • Set up the right cross-validation strategy
  • Establish a trivial baseline first
  • Compare 3–4 algorithm families with default params
  • Coarse random search → fine Bayesian search
  • Ensemble top models if complexity is acceptable
  • Evaluate on test set exactly once at the very end
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Key Takeaways
Summary

Key Takeaways

  • No free lunch — compare algorithms empirically on your data
  • Choose CV strategy by data type: stratified, time-series, or group
  • Grid search for small spaces; random search for many parameters
  • Bayesian optimization when evaluations are expensive
  • Ensembles (voting, stacking, blending) almost always improve performance
  • Evaluate on the test set exactly once — any further changes invalidate the estimate
13 / 13