Home / ML 101 / Module 10 / Lesson 2

Model Selection & Tuning

Choosing the right model and finding its best hyperparameters — the systematic science behind squeezing real-world performance.

~15 min read M10 · L2 Intermediate

No Free Lunch

A foundational result in ML theory — the No Free Lunch (NFL) theorem — states that no single algorithm outperforms all others across every possible problem. Averaged over all possible data distributions, every learning algorithm has the same expected performance. The implication is not pessimistic: it tells us that model selection must be informed by the specific structure of your data and task.

In practice, a small family of algorithms — gradient boosted trees, well-tuned neural networks, and ensembles — wins on most structured/tabular benchmarks. But "which model" always depends on your data size, feature types, latency constraints, and interpretability requirements. Empirical comparison on your specific problem is the only reliable guide.

The Model Selection Checklist

Before tuning: establish a strong baseline (e.g., mean prediction for regression, majority class for classification). Any model must beat the baseline to be worth deploying. Then compare 2–3 algorithm families before diving into hyperparameter tuning of any single model.

Cross-Validation

The bedrock of honest model evaluation. A single train/test split is fragile — performance can be high or low by chance depending on which samples land in each set. Cross-validation uses the data more fully by rotating which portion serves as the held-out test set.

K-Fold Cross-Validation

Split the data into K equal folds. Train on K−1 folds, evaluate on the remaining fold. Rotate through all K possibilities. Average the K scores. The result is a more stable estimate of generalization performance than a single split.

K-Fold CV Score
\text{CV Score} = \frac{1}{K}\sum_{k=1}^{K} \text{Score}_k
Average the evaluation metric over all K folds. Common choice: K = 5 or K = 10. Larger K gives lower bias but higher variance and compute cost.

Stratified K-Fold

For classification, standard K-fold may produce folds with very different class distributions by chance. Stratified K-fold ensures each fold has approximately the same class proportion as the full dataset. Always use stratification for imbalanced classification problems.

Cross-Validation Strategies by Data Type

Standard K-Fold
IID tabular data
Default choice when rows are independent. K=5 or K=10. Stratify for classification.
Time Series Split
Sequential data
Expanding window: train on past, validate on future. Never shuffle — future data cannot inform the past.
Group K-Fold
Grouped observations
When rows share an entity (same user, patient, store), keep all rows from one group in the same fold to prevent leakage.
Leave-One-Out
Very small datasets
K = N (each sample is its own fold). Maximizes training data but is computationally expensive.

Hyperparameter Search

Model hyperparameters — learning rate, tree depth, regularization strength, number of layers — are not learned from data; they are set before training. Finding good values is a search problem in a high-dimensional space.

Grid Search

Define a discrete grid of candidate values for each hyperparameter and evaluate every combination using cross-validation. Simple, exhaustive, and guaranteed to find the best point on the grid — but the cost grows exponentially with the number of hyperparameters.

Grid search is practical when you have ≤3 hyperparameters and a small candidate set per parameter. For a learning rate in {0.001, 0.01, 0.1}, max_depth in {3, 5, 7}, and n_estimators in {100, 300, 500}, grid search evaluates 3 × 3 × 3 = 27 configurations — feasible. Add two more parameters and it can become thousands.

Random Search

Instead of exhaustively trying every combination, sample configurations randomly from the hyperparameter space. Counterintuitively, random search often finds better configurations than grid search for the same compute budget. The reason: most hyperparameter spaces have low effective dimensionality — only a few parameters matter greatly, and random search explores more values of those key parameters.

Why Random Search Wins
\text{Distinct values per axis} = \begin{cases} n^{1/d} & \text{grid} \\ n & \text{random} \end{cases}
With n trials and d hyperparameters, random search projects n distinct values onto each axis. Grid search with the same budget uses only ⌊n^(1/d)⌋ distinct values per axis — far fewer when d is large.

Bayesian Optimization

Both grid and random search are uninformed — each trial ignores what previous trials revealed. Bayesian optimization builds a probabilistic surrogate model (usually a Gaussian process or tree-structured Parzen estimator) of the objective function and uses it to select the next configuration that balances exploitation (try near the current best) and exploration (try uncertain regions).

This is dramatically more sample-efficient than random search when each evaluation is expensive (e.g., training a deep neural network takes hours). Libraries: Optuna, Hyperopt, scikit-optimize. Rule of thumb: use Bayesian optimization when you have ≤50 trials to spend.

MethodBest forKey advantageKey limitation
Grid Search ≤3 params, small grids, reproducibility Exhaustive — finds best on the grid Exponential cost with dimensionality
Random Search Many params, limited budget Better coverage of important dimensions Uninformed — ignores prior results
Bayesian Opt. Expensive evaluations, limited trials Most sample-efficient Overhead of surrogate model fitting

Ensemble Methods

Model selection does not have to end with a single winner. Ensembles combine predictions from multiple models and almost always outperform any individual model — at the cost of added complexity and inference time.

Voting and Averaging

The simplest ensemble: train several diverse models independently and average their predictions (regression) or take a majority vote (classification). Works best when the base models make different kinds of errors — diversity is the key ingredient. Use models trained with different algorithms, different subsets of features, or different random seeds.

Stacking (Stacked Generalization)

Train a set of base learners on the training data. Then train a meta-learner (often a simple model like logistic regression or a linear model) on the out-of-fold predictions of the base learners. The meta-learner learns how to optimally combine the base predictions.

Stacking Anti-Leakage Rule

Never feed the meta-learner predictions from base models that have seen the same rows during training. Use out-of-fold (OOF) predictions: for each fold, train the base model on the other folds, then predict on the held-out fold. This gives a clean set of predictions the meta-learner can safely learn from.

Blending

A simpler variant of stacking: hold out a fixed blend set (say 20% of training data), train base models on the remaining 80%, predict on the blend set, and fit the meta-learner on those predictions. Less computationally expensive than stacking but wastes some training data and can be more sensitive to the blend set composition.

Voting / Averaging
Simplest ensemble
Average predictions or take majority vote. Diversity between base models is essential.
Stacking
Learned combination
Meta-learner trained on OOF predictions. More powerful but more complex.
Blending
Simpler stacking
Meta-learner trained on a fixed holdout set. Faster but wastes data.
Bagging
Variance reduction
Bootstrap samples + averaging. Random forests are the canonical example.

AutoML

Automated Machine Learning (AutoML) tools automate the full pipeline: feature preprocessing, model selection, hyperparameter tuning, and ensembling. They democratize ML by reducing the required expertise — but understanding what they do is still essential for diagnosing failures and deploying reliably.

Popular tools and their focus areas:

When to Use AutoML

AutoML excels at quickly establishing a strong baseline and at finding good pipelines when domain expertise is limited. It struggles at understanding business constraints (latency, model size, interpretability) and at debugging data quality issues. Use it as a starting point and a benchmark — not as a replacement for understanding your data.

Putting It Together: A Tuning Workflow

A reliable model selection and tuning process:

  1. Define your metric — what exactly are you optimizing? Accuracy? AUC-ROC? RMSE? Business metric? Misaligned metrics lead to misaligned models.
  2. Set up cross-validation — choose the right CV strategy for your data type. Never use the test set during search.
  3. Baseline — trivial predictor (mean, majority class). Any model must beat this.
  4. Compare algorithm families — run default hyperparameters for 3–4 candidate algorithms (linear model, gradient boosting, neural net). Pick the top 1–2.
  5. Coarse search — random search with wide ranges to identify the most important parameters and their rough optimal regions.
  6. Fine search — grid search or Bayesian optimization in the narrowed region.
  7. Ensemble — combine top models if the added complexity is worth it for your use case.
  8. Final evaluation — assess on the test set exactly once. Any further changes invalidate this estimate.

Key Takeaways

No single algorithm wins on all problems — always compare empirically. Use cross-validation appropriate to your data type (stratified for classification, time-series split for sequences, group split for grouped data). For hyperparameter search: grid search for small spaces, random search for many parameters, Bayesian optimization when evaluations are expensive. Ensembles (voting, stacking, blending) almost always improve performance at the cost of complexity. AutoML is a powerful baseline tool — understand what it does before trusting it.

Previous M10-L1: Feature Engineering Module Overview Next Lesson M10-L3: ML Pipeline