ML 101
M13 · L01
Module 13

Evolution as a Search Strategy

Every method so far asked the objective which way is downhill. This one never asks. It keeps a population of candidates and improves them by comparison alone.

01 / 13
ML 101
M13 · L01
The Gap

When There Is No Gradient

Gradient descent and backprop both need a smooth objective. Four common situations do not provide one.

  • Discrete choices — include feature 7, or not
  • Combinatorial — the order to visit 20 sites
  • Simulation-based — a score you must run, not solve
  • Black box — a vendor binary, or hardware
02 / 13
ML 101
M13 · L01
Worked Example

Which Features Should It See?

Eight candidate features means 256 subsets — try them all. Forty candidates means 2⁴⁰ subsets. At one training run per second that is roughly thirty-four thousand years.

8 features
256
40 features
1.1 × 10¹²
03 / 13
ML 101
M13 · L01
The Problem

Evaluate, Then Compare

Nothing here assumes the space is continuous, or that the objective has a derivative, or even that it is written down. The one operation assumed is evaluating it at a point you choose.

Black-Box Optimisation
x^{\star} = \arg\max_{x \in \mathcal{S}} f(x)
04 / 13
ML 101
M13 · L01
The Shift

One Point, or Many

Gradient descent moves a single point using a slope. A population never asks which direction is better — only which candidate is better.

What The Population Buys
Several regions searched at once, and parts of two parents combined into one child
05 / 13
ML 101
M13 · L01
The Bill

Population Is Not Free Power

One gradient step costs one forward and one backward pass. One generation costs one full evaluation per individual — and if an evaluation means training a model, that is a training run each.

Population
50
100 gens
5,000 evals
06 / 13
ML 101
M13 · L01
Holland, 1975

Four Ingredients

John Holland set out the scheme in Adaptation in Natural and Artificial Systems (1975); Goldberg’s 1989 textbook carried it into engineering.

  • Representation — how one candidate is written down
  • Fitness — the single number that scores it
  • Selection — who gets to be a parent
  • Variation — crossover and mutation
07 / 13
ML 101
M13 · L01
One Generation

Two Opposing Pressures

Selection is the only step that reads the fitness function, and it removes diversity. Variation creates it. The behaviour of the whole method is the balance between them.

Selection Then Variation
P_{t+1} = \mathrm{var}\big(\mathrm{sel}(P_t, f)\big)
08 / 13
ML 101
M13 · L01
Selection Pressure

Scale, Not Just Ranking

Scores of 10 and 9 give the winner 0.53 of the parent slots. Add 990 to both — nothing about which is better changes — and its share falls to 0.5003.

Fitness-Proportionate Selection
p_i = \frac{f(x_i)}{\sum_{j=1}^{N} f(x_j)}
09 / 13
ML 101
M13 · L01
The Classic Failure

Too Much Pressure

  • The best individual takes most of the parent slots
  • Its copies fill the population in a few generations
  • Crossover of identical parents returns the parent
  • Only mutation is left: a slow random walk
  • And it looks converged rather than stuck
10 / 13
ML 101
M13 · L01
The Family

Where the GA Sits

  • Random search — no memory, but a real baseline
  • Grid search — systematic, and multiplies per dimension
  • Hill climbing — one point, first local optimum wins
  • Simulated annealing — one point, sometimes accepts worse
  • Genetic algorithm — a population, plus recombination
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

A Strategy, Not a Model

A genetic algorithm produces no model. It chooses — features, hyperparameters, architectures — and your fitness function does the training. If the objective is differentiable, use the gradient.

13 / 13