From Theory to Code
The previous two lessons built up the theory of Support Vector Machines: maximum-margin classifiers, soft-margin slack variables, the kernel trick, and the dual formulation. Now it is time to use SVMs on real data. In this lesson we work through the full practical workflow: preprocessing, fitting, hyperparameter tuning via grid search, cross-validation, and evaluation.
We will use scikit-learn, Python’s standard machine learning library, which implements SVMs via the SVC (kernel SVM) and LinearSVC (linear SVM optimised for large datasets) classes. Both wrap the LIBSVM and LIBLINEAR solvers respectively.
Why Preprocessing Matters for SVMs
SVMs are not scale-invariant. The margin and kernel computations depend on Euclidean distances. A feature measured in kilometres will dominate a feature measured in metres, biasing the decision boundary. Always standardise your features before training an SVM.
The standard recipe is zero-mean, unit-variance scaling using StandardScaler:
from sklearn.preprocessing import StandardScaler from sklearn.pipeline import Pipeline from sklearn.svm import SVC # Always scale inside a pipeline to avoid data leakage pipe = Pipeline([ ('scaler', StandardScaler()), ('svm', SVC(kernel='rbf', C=1.0, gamma='scale')) ])
Using a Pipeline is critical: it ensures that the scaler is fit only on training data and applied consistently to validation and test folds. Fitting the scaler on the entire dataset before splitting — a common mistake — leaks statistics from the test set into training, producing over-optimistic accuracy estimates.
Data leakage is one of the most common sources of over-optimistic results in ML. When you scale using the full dataset before splitting, your test set statistics influence the scaler, which then influences training. The Pipeline abstraction prevents this by ensuring preprocessing is always refitted within each cross-validation fold.
The SVC API
The key parameters of SVC are:
| Parameter | Default | Meaning |
|---|---|---|
| kernel | 'rbf' |
Kernel type: 'linear', 'poly', 'rbf', 'sigmoid', or a callable |
| C | 1.0 |
Regularisation strength. Smaller C → wider margin, more misclassifications allowed |
| gamma | 'scale' |
RBF/poly/sigmoid bandwidth. 'scale' = 1/(n_features · X.var()). Also accepts a float. |
| degree | 3 |
Degree for polynomial kernel only |
| probability | False |
Enable probability estimates via Platt scaling (slower) |
| class_weight | None |
Set to 'balanced' for imbalanced classes |
Multi-Class Classification
An SVM is fundamentally a binary classifier — it finds a single hyperplane separating two classes. To handle K > 2 classes, the problem is decomposed into many binary sub-problems using one of two schemes.
One-vs-one (OvO) trains one classifier for every pair of classes — K(K−1)/2 of them. At prediction time each classifier casts a vote for one of its two classes, and the class with the most votes wins. This is what SVC does internally: LIBSVM is always one-vs-one, and the decision_function_shape parameter only reshapes the output scores — it does not change the training scheme. Each sub-problem sees only the two relevant classes, so individual fits are cheap, but the classifier count grows quadratically with K.
One-vs-rest (OvR) — also called one-vs-all — trains one classifier per class, K in total, each separating its class from all the others combined. The class whose decision function scores highest wins. LinearSVC uses OvR by default. It needs only K classifiers (linear in the number of classes), but each is trained on the entire dataset.
One-vs-one builds O(K²) classifiers, each on a small two-class subset; one-vs-rest builds O(K) classifiers, each on the full data. For a handful of classes OvO is common (and the default for kernel SVC); when K is large, OvR keeps the number of models manageable. Either way, scikit-learn performs the decomposition for you — you still call fit and predict once.
Cross-Validation Strategy
A single train/test split gives a noisy estimate of generalisation performance. k-fold cross-validation (typically k = 5 or 10) partitions the data into k folds, trains on k−1 folds, and evaluates on the held-out fold, rotating through all k possibilities. The reported score is the mean (and optionally standard deviation) across folds.
from sklearn.model_selection import cross_val_score, StratifiedKFold cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42) scores = cross_val_score(pipe, X_train, y_train, cv=cv, scoring='accuracy') print(f"CV accuracy: {scores.mean():.3f} ± {scores.std():.3f}")
StratifiedKFold preserves class proportions in each fold, which matters for imbalanced datasets. Shuffling before splitting prevents fold assignments that depend on data order (e.g., if data was sorted by class).
Hyperparameter Search: C and γ
The two most important hyperparameters for an RBF SVM are C and γ. They interact: the optimal C depends on γ and vice versa. This means you must search over a grid of (C, γ) pairs, not tune each independently.
Search on a logarithmic scale: the SVM response to C and γ is roughly log-linear, so testing values like {0.01, 0.1, 1, 10, 100} is more informative than a linear grid {1, 2, 3, 4, 5}.
from sklearn.model_selection import GridSearchCV import numpy as np param_grid = { 'svm__C': np.logspace(-2, 3, 6), # 0.01 to 1000 'svm__gamma': np.logspace(-4, 1, 6), # 0.0001 to 10 } search = GridSearchCV( pipe, param_grid, cv=5, scoring='f1_weighted', n_jobs=-1, verbose=1 ) search.fit(X_train, y_train) print(search.best_params_) print(f"Best CV F1: {search.best_score_:.3f}")
Note the double-underscore syntax 'svm__C': this is how GridSearchCV passes parameters to a named step inside a Pipeline. The step name is the key you gave in the Pipeline constructor ('svm' here), and the parameter name follows the double underscore.
If your grid is large (many parameters, wide ranges), RandomizedSearchCV is more efficient than exhaustive grid search. It samples a fixed number of parameter combinations from the specified distributions, often finding a near-optimal solution in a fraction of the time. Use it when grid search would take hours.
Evaluating SVM Performance
Accuracy is a misleading metric when classes are imbalanced. A classifier that always predicts the majority class can achieve 90% accuracy on a 90/10 split while being completely useless for the minority class. Use precision, recall, F1, and the confusion matrix instead.
from sklearn.metrics import classification_report, confusion_matrix best_model = search.best_estimator_ y_pred = best_model.predict(X_test) print(classification_report(y_test, y_pred)) print(confusion_matrix(y_test, y_pred))
The classification_report shows per-class precision, recall, and F1 along with macro/weighted averages. For binary problems, choose a scoring metric aligned with your deployment goal: precision if false positives are costly (spam filter), recall if false negatives are costly (disease screening), F1 if you need a balance.
Decision Scores and Calibrated Probabilities
By default, SVC.predict() returns a hard class label. You can get the raw decision function value (signed distance to the margin) via decision_function(). For probability estimates, set probability=True in the constructor: scikit-learn uses Platt scaling — fitting a logistic regression on top of the SVM outputs using an internal cross-validation — to convert decision scores to probabilities.
Platt scaling increases training time (because it requires an internal cross-validation) and may not perfectly calibrate probabilities, especially in the tails. For ranking tasks (where you need ordering of predictions, not calibrated probabilities), decision_function() is usually sufficient.
LinearSVC for Large Datasets
For large datasets (>100,000 examples) with a linear kernel, LinearSVC is dramatically faster than SVC(kernel='linear'). It uses the LIBLINEAR solver (primal coordinate descent) instead of LIBSVM (SMO on the dual), which scales as O(n) in training examples rather than O(n2) to O(n3).
from sklearn.svm import LinearSVC # LinearSVC: fast linear SVM, no kernel trick, no predict_proba pipe_linear = Pipeline([ ('scaler', StandardScaler()), ('svm', LinearSVC(C=1.0, max_iter=2000)) ])
Trade-offs of LinearSVC: it does not support kernel functions (linear only), does not directly output probabilities, and has slightly different regularisation conventions than SVC. For text classification, genomics, and other high-dimensional sparse problems, it is often the best choice.
Support Vector Regression (SVR)
The same maximum-margin machinery solves regression, not just classification. Support Vector Regression fits a function that stays as flat as possible while keeping most training points within a margin of tolerance around it.
The key idea is the ε-insensitive tube: errors smaller than ε are ignored entirely — any prediction within ±ε of the true value incurs zero loss. Only points lying on or outside the tube become support vectors and shape the fit. This is the mirror image of classification, where the support vectors are the points on or inside the margin.
Because most points typically fall inside the tube, the solution is sparse — determined by relatively few support vectors, exactly as in classification. SVR reuses the identical kernel arsenal (linear, polynomial, RBF), so a kernel SVR captures non-linear relationships the same way a kernel SVC captures non-linear boundaries. scikit-learn exposes it as SVR, NuSVR, and the fast linear-only LinearSVR, mirroring the classifier classes. Two knobs govern it: ε sets the tube width, and C trades off tube violations against flatness — the same role C plays in SVC.
Practical Tips and Common Pitfalls
Start with a linear SVM. It is fast, interpretable, and often competitive on high-dimensional data. Only move to RBF if the linear kernel underfits.
Always scale. Forgetting StandardScaler is the single most common SVM mistake. Even with gamma='scale' (which normalises by feature variance), mean-centering via the full scaler helps.
Grid search on a log scale. Searching C in {1, 2, 3} while γ varies over orders of magnitude will miss the optimum. Use np.logspace().
Watch for convergence warnings. If LinearSVC or SVC prints a convergence warning, increase max_iter. Do not ignore warnings — the model may not have converged to a valid solution.
Imbalanced classes. Set class_weight='balanced' when one class is much rarer than the other. This adjusts the penalty C per class proportionally to class frequency, giving the minority class more influence on the margin.
Support vector count as a sanity check. After fitting, inspect model.n_support_. If almost every training point is a support vector, the model is likely underfitting (C too small or γ too large in the wrong direction). If very few are support vectors, it may be overfitting (C too large).
SVMs shine on small to medium, high-dimensional datasets where the number of features is comparable to or exceeds the number of examples — text classification, bioinformatics, image feature vectors. They also work well when the classes are nearly linearly separable (linear SVM) or when you need a non-parametric decision boundary (RBF SVM) on clean, well-preprocessed data. For very large datasets, tree-based ensembles (gradient boosting) or neural networks usually outperform SVMs because they scale more favourably.
- Always standardise features before training an SVM — use
StandardScalerinside aPipelineto prevent data leakage. - Key
SVCparameters:kernel(default RBF),C(regularisation),gamma(RBF bandwidth). - Use
StratifiedKFoldcross-validation to get a reliable, unbiased estimate of generalisation performance. - Tune C and γ jointly via grid search on a log scale — they interact and cannot be tuned independently.
- For imbalanced datasets, use
class_weight='balanced'and evaluate with F1 or AUC rather than accuracy. - For large datasets with a linear kernel, prefer
LinearSVC— it scales as O(n) instead of O(n²–n³). - SVM is binary at heart: multi-class uses one-vs-one (
SVC, O(K²) classifiers) or one-vs-rest (LinearSVC, O(K)); SVR extends the same margin idea to regression via an ε-insensitive tube. - SVMs excel on small-to-medium, high-dimensional, well-preprocessed datasets; tree ensembles and neural networks tend to outperform on very large datasets.