When Evolution Wins
And when it loses. Where a genetic algorithm earns its keep, where a cheaper method already exists, and the cost argument that decides most of it.
Generality Is Not Free
A genetic algorithm asks almost nothing of a problem: no derivative, no convexity, not even numbers. That is exactly why it is slow — you pay for every assumption you decline to make.
Use the Gradient
One backward pass returns the exact derivative for every parameter. P fitness evaluations return P scalars, from which a direction in d dimensions must be guessed.
The Multiplication You Commit To
Population times generations evaluations, against one forward and one backward pass per gradient step. Reach for evolution when that trade is forced by the problem, not chosen.
A Genuine Black Box
- Not differentiable in any usable way
- Continuous, discrete and conditional choices mixed together
- Every evaluation is a full training run
- So evolution is admissible here — which is not the same as best
Four Ways to Search
- Grid — enumerates a grid fixed before the first run. Fully parallel
- Random — samples ranges, ignores every result. Fully parallel
- Bayesian — fits a surrogate to all results, then optimises an acquisition function. Sequential by design
- Genetic — recombines the better members of a population. Parallel within a generation
Which Structure Fits
- Bayesian search extracts more from each evaluation — rank selection throws magnitudes away
- A genetic algorithm fills parallel workers with no extra theory
- Awkward spaces — variable length, conditional, permutations — favour evolution
- Baseline first: random search over sensible ranges (Bergstra & Bengio, 2012)
Feature Selection
One bit per feature, one string per subset. Fitness is a cross-validated score with a penalty for keeping features — a wrapper method, so it can find features that only help together.
Overfitting the Validation Score
You are selecting the maximum of thousands of noisy scores, so the winner looks better than it is. Hold out a test set the search never touches.
- Rule out L1 first — sparsity for the price of one fit (Lesson 2.3)
- Tree and permutation importance are nearly free (Lesson 3.2)
- Greedy forward or backward selection costs far fewer evaluations
Weights or Architecture
- Weights — wrong choice whenever a differentiable loss exists
- Reasonable when the score comes from a simulation, a game, or a delayed reward
- Architecture — depth, width and connectivity are discrete. No gradient exists, so search is honest
- Module 14 attacks the same class with reinforcement learning
NEAT in Outline
- Start minimal, complexify — add nodes and connections over the run
- Historical markings — so crossover aligns genes with shared ancestry
- Speciation — protects a new structure long enough for its weights to be optimised
Genetic Programming
The individual is a program, held as an expression tree. Crossover swaps subtrees, so the child is still a valid program. Applied to data, this is symbolic regression.
When It Loses
- The objective is differentiable — use the gradient
- Convex, linear or integer program — a solver returns a bound; a genetic algorithm returns none
- Known combinatorial structure, or a space small enough to enumerate
- Fitness so noisy that selection cannot tell candidates apart
- A fitness function you do not trust — that is Lesson 13.3, not the algorithm
Key Takeaways
- Generality costs evaluations — differentiable means use the gradient
- Hyperparameters: compare with grid, random and Bayesian search before committing
- Feature selection works as subset search, and overfits validation if unguarded
- Evolving architecture is more defensible than evolving weights
- Genetic programming buys a readable formula, at the price of bloat
- Up next: Module 14 turns to learning from reward