The Kernel Trick
When classes cannot be separated by a line, map them to a space where they can — but never compute that mapping. This elegant shortcut is the kernel trick.
When Lines Fail
Many real patterns — concentric rings, XOR, spirals — cannot be separated by any hyperplane. A linear SVM would fail completely on these.
Lifting to Higher Dimensions
A feature map φ transforms data from the original space into a higher-dimensional space where the classes become linearly separable.
Kernel Function
In the SVM dual, data appears only as dot products φ(x)·φ(z). A kernel function computes this inner product without ever building φ(x).
Mercer’s Condition
- A valid kernel must have a positive semi-definite kernel matrix
- Guarantees K(x,z) computes a true inner product somewhere
- Standard kernels (linear, RBF, polynomial) all satisfy it
- Custom kernels must be verified
- Violation → optimization may not converge
Common Kernels
RBF / Gaussian Kernel
Similarity decreases as a Gaussian with distance. Corresponds to an infinite-dimensional feature space — yet tractable. A universal approximator.
The γ Parameter
- Large γ — narrow Gaussian → only very close points similar → wiggly boundary
- Small γ — wide Gaussian → distant points still similar → smoother boundary
- Large γ risks overfitting; small γ risks underfitting
- Default: γ = 1 / n_features
- Tune jointly with C via grid search
Lagrange Multipliers
The SVM dual replaces the primal with Lagrange multipliers αi. Data appears only as dot products — so any kernel can be plugged in.
Prediction via Kernels
After training, prediction is a weighted sum of kernels against the support vectors only (αi > 0). Fast even in infinite-dimensional spaces.
Scalability Challenge
- Kernel matrix is n × n — O(n²) memory
- Solving the dual: O(n³) time
- Practical limit: ~100k training examples
- Workarounds: Nyström approximation, random features
- Deep learning scales better for very large data
- Kernel SVM wins on small to medium datasets
What You Learned
Kernels let SVMs find nonlinear boundaries by implicitly mapping to high-dimensional spaces. K(x,z) = φ(x)·φ(z) without ever computing φ. RBF is the default choice; γ and C both need tuning. Next: SVM in Practice — training, tuning, and deployment.