ML 101
M13 · L04
Evolutionary Search

When Evolution Wins

And when it loses. Where a genetic algorithm earns its keep, where a cheaper method already exists, and the cost argument that decides most of it.

01 / 15
ML 101
M13 · L04
The Right Question

Generality Is Not Free

A genetic algorithm asks almost nothing of a problem: no derivative, no convexity, not even numbers. That is exactly why it is slow — you pay for every assumption you decline to make.

Reframe It
Never "is a genetic algorithm good?" Always "does my problem have structure a cheaper method could exploit?"
02 / 15
ML 101
M13 · L04
The First Rule

Use the Gradient

One backward pass returns the exact derivative for every parameter. P fitness evaluations return P scalars, from which a direction in d dimensions must be guessed.

What Backprop Returns
\nabla_{\boldsymbol{\theta}} L = \left[\frac{\partial L}{\partial \theta_1}, \ldots, \frac{\partial L}{\partial \theta_d}\right]
03 / 15
ML 101
M13 · L04
Cost Accounting

The Multiplication You Commit To

Population times generations evaluations, against one forward and one backward pass per gradient step. Reach for evolution when that trade is forced by the problem, not chosen.

Evaluation Budget
N_{\text{eval}} = P \cdot G
04 / 15
ML 101
M13 · L04
Hyperparameters

A Genuine Black Box

  • Not differentiable in any usable way
  • Continuous, discrete and conditional choices mixed together
  • Every evaluation is a full training run
  • So evolution is admissible here — which is not the same as best
05 / 15
ML 101
M13 · L04
Compared Honestly

Four Ways to Search

  • Grid — enumerates a grid fixed before the first run. Fully parallel
  • Random — samples ranges, ignores every result. Fully parallel
  • Bayesian — fits a surrogate to all results, then optimises an acquisition function. Sequential by design
  • Genetic — recombines the better members of a population. Parallel within a generation
06 / 15
ML 101
M13 · L04
Mechanism, Not a Winner

Which Structure Fits

  • Bayesian search extracts more from each evaluation — rank selection throws magnitudes away
  • A genetic algorithm fills parallel workers with no extra theory
  • Awkward spaces — variable length, conditional, permutations — favour evolution
  • Baseline first: random search over sensible ranges (Bergstra & Bengio, 2012)
07 / 15
ML 101
M13 · L04
Subset Search

Feature Selection

One bit per feature, one string per subset. Fitness is a cross-validated score with a penalty for keeping features — a wrapper method, so it can find features that only help together.

Space and Fitness
|\mathcal{S}| = 2^{d}, \quad F(S) = \mathrm{CV}(S) - \alpha \frac{|S|}{d}
08 / 15
ML 101
M13 · L04
The Trap

Overfitting the Validation Score

You are selecting the maximum of thousands of noisy scores, so the winner looks better than it is. Hold out a test set the search never touches.

  • Rule out L1 first — sparsity for the price of one fit (Lesson 2.3)
  • Tree and permutation importance are nearly free (Lesson 3.2)
  • Greedy forward or backward selection costs far fewer evaluations
09 / 15
ML 101
M13 · L04
Neuroevolution

Weights or Architecture

  • Weights — wrong choice whenever a differentiable loss exists
  • Reasonable when the score comes from a simulation, a game, or a delayed reward
  • Architecture — depth, width and connectivity are discrete. No gradient exists, so search is honest
  • Module 14 attacks the same class with reinforcement learning
10 / 15
ML 101
M13 · L04
Stanley & Miikkulainen (2002)

NEAT in Outline

  • Start minimal, complexify — add nodes and connections over the run
  • Historical markings — so crossover aligns genes with shared ancestry
  • Speciation — protects a new structure long enough for its weights to be optimised
11 / 15
ML 101
M13 · L04
Koza (1992)

Genetic Programming

The individual is a program, held as an expression tree. Crossover swaps subtrees, so the child is still a valid program. Applied to data, this is symbolic regression.

The Deliverable, and the Hazard
You get a readable formula rather than weights. Trees grow without fitting better — bloat — so charge for size.
12 / 15
ML 101
M13 · L04
Do Not Reach For It

When It Loses

  • The objective is differentiable — use the gradient
  • Convex, linear or integer program — a solver returns a bound; a genetic algorithm returns none
  • Known combinatorial structure, or a space small enough to enumerate
  • Fitness so noisy that selection cannot tell candidates apart
  • A fitness function you do not trust — that is Lesson 13.3, not the algorithm
13 / 15
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

14 / 15
ML 101
Key Takeaways
Module 13 Complete

Key Takeaways

  • Generality costs evaluations — differentiable means use the gradient
  • Hyperparameters: compare with grid, random and Bayesian search before committing
  • Feature selection works as subset search, and overfits validation if unguarded
  • Evolving architecture is more defensible than evolving weights
  • Genetic programming buys a readable formula, at the price of bloat
  • Up next: Module 14 turns to learning from reward
15 / 15