SVM in Practice
Theory is only half the story. Now we train, tune, and evaluate real SVMs using scikit-learn — from raw data to a deployed classifier.
Always Scale
SVMs are not scale-invariant — distances drive the margin. A feature in km dominates one in mm. Always standardise inside a Pipeline to prevent data leakage.
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
pipe = Pipeline([('sc', StandardScaler()), ('svm', SVC())])
The SVC API
- kernel —
'rbf'(default),'linear','poly' - C — regularisation strength (default 1.0)
- gamma — RBF bandwidth (
'scale'default) - probability — enable Platt scaling (slower)
- class_weight —
'balanced'for imbalanced data
Cross-Validation
A single split gives noisy performance estimates. k-fold CV averages over k held-out folds — each data point is tested exactly once.
Stratified Folds
- Use StratifiedKFold to preserve class proportions
- Shuffle before splitting (set random_state)
- 5 or 10 folds is standard
- Fit preprocessing inside each fold via Pipeline
- Report mean ± std across folds
Grid Search
C and γ interact — tune them jointly, not independently. Search on a log scale: {0.01, 0.1, 1, 10, 100}. Use GridSearchCV with your pipeline.
C & γ Trade-offs
- Large C — small margin, few errors allowed → risk overfitting
- Small C — wide margin, more errors allowed → risk underfitting
- Large γ — narrow RBF bell → wiggly boundary
- Small γ — wide RBF bell → smooth boundary
- High C + high γ = overfitting. Low C + low γ = underfitting
Evaluation Metrics
- Precision — of predicted positives, how many are correct?
- Recall — of actual positives, how many did we catch?
- F1 — harmonic mean of precision and recall
- Confusion matrix — full breakdown of TP/FP/TN/FN
- Accuracy misleads on imbalanced classes
Platt Scaling
Set probability=True to get calibrated probabilities. Platt scaling fits a logistic sigmoid on top of the SVM decision scores.
Beyond Binary
- Multi-class — SVM is binary;
SVCfits one-vs-one, K(K−1)/2 models, majority vote - One-vs-rest —
LinearSVCdefault: K models, one per class, argmax wins - SVR — regression via an ε-insensitive tube; errors inside ±ε cost nothing
- Support vectors = points on or outside the tube; same kernels & C as SVC
- LinearSVC scales to millions (LIBLINEAR, O(n)) — text, genomics, sparse data
Common Mistakes
- Scaling outside the Pipeline → data leakage
- Linear search for C and γ → miss optimum
- Ignoring convergence warnings → invalid model
- Using accuracy on imbalanced data → misleading
- Tuning C and γ independently → suboptimal
- Forgetting class_weight on skewed classes
What You Learned
Scale features inside a Pipeline. Tune C & γ jointly via log-scale grid search. Use StratifiedKFold CV. Evaluate with F1 and confusion matrix. For large data, use LinearSVC. Module 4 complete!