Maximum Margin
Among all lines that separate two classes, one is uniquely best: the one that is as far as possible from both groups. This is the core idea of Support Vector Machines.
Linear Separability
Two classes are linearly separable if there exists at least one hyperplane that perfectly separates them. In 2D this is a line; in 3D, a plane; in higher dimensions, a hyperplane.
The Hyperplane
The boundary is defined by weight vector w and bias b. Points on opposite sides get opposite signs — that determines the predicted class.
Maximizing the Margin
Two parallel margin planes flank the boundary at distance 1/‖w‖ each. The total margin is 2/‖w‖. To maximize margin, minimize ‖w‖.
Support Vectors
- Points that lie exactly on the margin planes
- Only they determine the decision boundary
- Remove any other point — boundary unchanged
- Sparse representation: only support vectors stored
- Typically a small fraction of training data
Hard Margin SVM
When data is perfectly separable, the hard margin SVM finds the widest margin with zero misclassification. It is a convex quadratic program with a unique global solution.
Soft Margin & Slack
Real data is rarely perfectly separable. Slack variables ξi allow some points to violate the margin. ξi = 0 is correct; 0<ξi≤1 is inside margin; ξi>1 is misclassified.
The C Parameter
- Large C — high penalty for errors → narrow margin, low bias
- Small C — low penalty → wide margin, more errors allowed
- C → ∞ approaches hard margin SVM
- Tune via cross-validation on log scale
- Typical range: 0.001 to 1000
Margin Diagram
The decision boundary (solid) is flanked by two margin planes (dashed). Support vectors sit exactly on the margin planes. Everything else is irrelevant to the boundary.
Feature Scaling
SVMs measure margin in the original feature space. A feature with large values dominates the weight norm. Always standardize features before training — zero mean, unit variance.
When Lines Fail
A linear hyperplane cannot separate classes that form circles, spirals, or other nonlinear structures. The kernel trick overcomes this by implicitly mapping data to a higher-dimensional space.
What You Learned
SVMs find the hyperplane with the widest margin between classes. Support vectors alone define the boundary. The C parameter trades off margin width vs. misclassification. Next: The Kernel Trick extends SVMs to nonlinear problems.