ML 101
M04 · L02
Module 4

The Kernel Trick

When classes cannot be separated by a line, map them to a space where they can — but never compute that mapping. This elegant shortcut is the kernel trick.

01 / 13
ML 101
M04 · L02
The Problem

When Lines Fail

Many real patterns — concentric rings, XOR, spirals — cannot be separated by any hyperplane. A linear SVM would fail completely on these.

XOR
Not separable
Circles
Not separable
02 / 13
ML 101
M04 · L02
Feature Mapping

Lifting to Higher Dimensions

A feature map φ transforms data from the original space into a higher-dimensional space where the classes become linearly separable.

Feature Map
\phi: \mathbb{R}^n \to \mathcal{F}
03 / 13
ML 101
M04 · L02
The Shortcut

Kernel Function

In the SVM dual, data appears only as dot products φ(x)·φ(z). A kernel function computes this inner product without ever building φ(x).

Kernel Trick
K(\mathbf{x}, \mathbf{z}) = \phi(\mathbf{x})^\top \phi(\mathbf{z})
04 / 13
ML 101
M04 · L02
Validity Condition

Mercer’s Condition

  • A valid kernel must have a positive semi-definite kernel matrix
  • Guarantees K(x,z) computes a true inner product somewhere
  • Standard kernels (linear, RBF, polynomial) all satisfy it
  • Custom kernels must be verified
  • Violation → optimization may not converge
05 / 13
ML 101
M04 · L02
Kernel Zoo

Common Kernels

Linear
K = x·z
Polynomial
K = (x·z + c)d
RBF
K = exp(−γ‖x−z‖²)
Sigmoid
K = tanh(κx·z + θ)
06 / 13
ML 101
M04 · L02
Most Popular

RBF / Gaussian Kernel

Similarity decreases as a Gaussian with distance. Corresponds to an infinite-dimensional feature space — yet tractable. A universal approximator.

RBF Kernel
K(\mathbf{x}, \mathbf{z}) = \exp\!\left(-\gamma \|\mathbf{x} - \mathbf{z}\|^2\right)
07 / 13
ML 101
M04 · L02
RBF Hyperparameter

The γ Parameter

  • Large γ — narrow Gaussian → only very close points similar → wiggly boundary
  • Small γ — wide Gaussian → distant points still similar → smoother boundary
  • Large γ risks overfitting; small γ risks underfitting
  • Default: γ = 1 / n_features
  • Tune jointly with C via grid search
08 / 13
ML 101
M04 · L02
Dual Formulation

Lagrange Multipliers

The SVM dual replaces the primal with Lagrange multipliers αi. Data appears only as dot products — so any kernel can be plugged in.

Dual Objective
\max_{\boldsymbol{\alpha}}\;\sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j K(\mathbf{x}_i,\mathbf{x}_j)
09 / 13
ML 101
M04 · L02
At Test Time

Prediction via Kernels

After training, prediction is a weighted sum of kernels against the support vectors only (αi > 0). Fast even in infinite-dimensional spaces.

Prediction
f(\mathbf{x}) = \text{sign}\!\left(\sum_{i:\alpha_i>0} \alpha_i y_i K(\mathbf{x}_i, \mathbf{x}) + b\right)
10 / 13
ML 101
M04 · L02
Practical Limits

Scalability Challenge

  • Kernel matrix is n × n — O(n²) memory
  • Solving the dual: O(n³) time
  • Practical limit: ~100k training examples
  • Workarounds: Nyström approximation, random features
  • Deep learning scales better for very large data
  • Kernel SVM wins on small to medium datasets
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

Kernels let SVMs find nonlinear boundaries by implicitly mapping to high-dimensional spaces. K(x,z) = φ(x)·φ(z) without ever computing φ. RBF is the default choice; γ and C both need tuning. Next: SVM in Practice — training, tuning, and deployment.

Next Lesson
13 / 13