Model Selection & Tuning
Choosing the right model and finding its best hyperparameters — the systematic science behind squeezing real-world performance out of your ML pipeline.
No Single Winner
Averaged over all possible problems, every algorithm performs equally. There is no universally best model — the right choice depends on your data, size, features, and constraints.
Cross-Validation
A single train/test split can be lucky or unlucky. K-fold cross-validation rotates the held-out set across K folds and averages the scores:
Match CV to Your Data Type
- Stratified K-Fold — preserves class ratios. Use for classification
- Time Series Split — expanding window: always train on past, validate on future
- Group K-Fold — keep all rows from one entity (user, patient) in the same fold
- Leave-One-Out — K = N; for very small datasets only
Grid Search vs. Random Search
- Grid search — tries every combination. Guaranteed best on grid, but exponential cost
- Random search — samples configurations randomly. Covers more values of important parameters for the same budget
- For ≤3 parameters with small grids: grid search. Otherwise: random search
- Bergstra & Bengio (2012): random search finds better results than grid search in practice
Coverage vs. Exhaustion
With n trials and d parameters, grid search projects only n^(1/d) distinct values per axis. Random search projects n distinct values per axis — far more when d is large.
Bayesian Optimization
Builds a surrogate model of the objective and picks the next configuration by balancing exploitation (near current best) and exploration (uncertain regions).
Use when evaluations are expensive (deep learning, long training). Tools: Optuna, Hyperopt, scikit-optimize.
Combine for Strength
- Voting / Averaging — train diverse models independently, average predictions
- Stacking — meta-learner trained on out-of-fold predictions of base models
- Blending — meta-learner trained on a fixed holdout set. Simpler but wastes data
- Key insight: diversity beats accuracy for individual base models
Out-of-Fold Predictions
Never feed the meta-learner predictions from base models trained on the same rows. Use out-of-fold (OOF) predictions: for each fold, train on others, predict on this fold.
Automated Pipelines
- Auto-sklearn — Bayesian opt over scikit-learn pipelines
- TPOT — genetic programming; exports clean code
- H2O AutoML — stacks GBMs and neural nets; production-ready
- AutoGluon — strong tabular defaults from AWS
- Use as a baseline and benchmark — understand what it does
A Systematic Process
- Define your metric — what exactly are you optimizing?
- Set up the right cross-validation strategy
- Establish a trivial baseline first
- Compare 3–4 algorithm families with default params
- Coarse random search → fine Bayesian search
- Ensemble top models if complexity is acceptable
- Evaluate on test set exactly once at the very end
Key Takeaways
- No free lunch — compare algorithms empirically on your data
- Choose CV strategy by data type: stratified, time-series, or group
- Grid search for small spaces; random search for many parameters
- Bayesian optimization when evaluations are expensive
- Ensembles (voting, stacking, blending) almost always improve performance
- Evaluate on the test set exactly once — any further changes invalidate the estimate