ML 101
M04 · L03
Module 4

SVM in Practice

Theory is only half the story. Now we train, tune, and evaluate real SVMs using scikit-learn — from raw data to a deployed classifier.

01 / 13
ML 101
M04 · L03
Step Zero

Always Scale

SVMs are not scale-invariant — distances drive the margin. A feature in km dominates one in mm. Always standardise inside a Pipeline to prevent data leakage.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
pipe = Pipeline([('sc', StandardScaler()), ('svm', SVC())])
02 / 13
ML 101
M04 · L03
Key Parameters

The SVC API

  • kernel — 'rbf' (default), 'linear', 'poly'
  • C — regularisation strength (default 1.0)
  • gamma — RBF bandwidth ('scale' default)
  • probability — enable Platt scaling (slower)
  • class_weight — 'balanced' for imbalanced data
03 / 13
ML 101
M04 · L03
Reliable Estimation

Cross-Validation

A single split gives noisy performance estimates. k-fold CV averages over k held-out folds — each data point is tested exactly once.

CV Score
\hat{S}_{\text{CV}} = \tfrac{1}{k}\sum_{i=1}^{k} S_i
04 / 13
ML 101
M04 · L03
Best Practice

Stratified Folds

  • Use StratifiedKFold to preserve class proportions
  • Shuffle before splitting (set random_state)
  • 5 or 10 folds is standard
  • Fit preprocessing inside each fold via Pipeline
  • Report mean ± std across folds
05 / 13
ML 101
M04 · L03
Tuning C & γ

Grid Search

C and γ interact — tune them jointly, not independently. Search on a log scale: {0.01, 0.1, 1, 10, 100}. Use GridSearchCV with your pipeline.

C range
10⁻² → 10²
γ range
10⁻⁴ → 10¹
06 / 13
ML 101
M04 · L03
Hyperparameter Effects

C & γ Trade-offs

  • Large C — small margin, few errors allowed → risk overfitting
  • Small C — wide margin, more errors allowed → risk underfitting
  • Large γ — narrow RBF bell → wiggly boundary
  • Small γ — wide RBF bell → smooth boundary
  • High C + high γ = overfitting. Low C + low γ = underfitting
07 / 13
ML 101
M04 · L03
Beyond Accuracy

Evaluation Metrics

  • Precision — of predicted positives, how many are correct?
  • Recall — of actual positives, how many did we catch?
  • F1 — harmonic mean of precision and recall
  • Confusion matrix — full breakdown of TP/FP/TN/FN
  • Accuracy misleads on imbalanced classes
08 / 13
ML 101
M04 · L03
Probability Output

Platt Scaling

Set probability=True to get calibrated probabilities. Platt scaling fits a logistic sigmoid on top of the SVM decision scores.

Platt
P(y{=}1\mid f) = \sigma(Af+B) = \dfrac{1}{1+e^{-(Af+B)}}
09 / 13
ML 101
M04 · L03
SVM Variants

Beyond Binary

  • Multi-class — SVM is binary; SVC fits one-vs-one, K(K−1)/2 models, majority vote
  • One-vs-rest — LinearSVC default: K models, one per class, argmax wins
  • SVR — regression via an ε-insensitive tube; errors inside ±ε cost nothing
  • Support vectors = points on or outside the tube; same kernels & C as SVC
  • LinearSVC scales to millions (LIBLINEAR, O(n)) — text, genomics, sparse data
10 / 13
ML 101
M04 · L03
Pitfalls

Common Mistakes

  • Scaling outside the Pipeline → data leakage
  • Linear search for C and γ → miss optimum
  • Ignoring convergence warnings → invalid model
  • Using accuracy on imbalanced data → misleading
  • Tuning C and γ independently → suboptimal
  • Forgetting class_weight on skewed classes
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

Scale features inside a Pipeline. Tune C & γ jointly via log-scale grid search. Use StratifiedKFold CV. Evaluate with F1 and confusion matrix. For large data, use LinearSVC. Module 4 complete!

Module 4 Done
SVMs: Theory & Practice →
13 / 13