Evolution as a Search Strategy
Every method so far asked the objective which way is downhill. This one never asks. It keeps a population of candidates and improves them by comparison alone.
When There Is No Gradient
Gradient descent and backprop both need a smooth objective. Four common situations do not provide one.
- Discrete choices — include feature 7, or not
- Combinatorial — the order to visit 20 sites
- Simulation-based — a score you must run, not solve
- Black box — a vendor binary, or hardware
Which Features Should It See?
Eight candidate features means 256 subsets — try them all. Forty candidates means 2⁴⁰ subsets. At one training run per second that is roughly thirty-four thousand years.
Evaluate, Then Compare
Nothing here assumes the space is continuous, or that the objective has a derivative, or even that it is written down. The one operation assumed is evaluating it at a point you choose.
One Point, or Many
Gradient descent moves a single point using a slope. A population never asks which direction is better — only which candidate is better.
Population Is Not Free Power
One gradient step costs one forward and one backward pass. One generation costs one full evaluation per individual — and if an evaluation means training a model, that is a training run each.
Four Ingredients
John Holland set out the scheme in Adaptation in Natural and Artificial Systems (1975); Goldberg’s 1989 textbook carried it into engineering.
- Representation — how one candidate is written down
- Fitness — the single number that scores it
- Selection — who gets to be a parent
- Variation — crossover and mutation
Two Opposing Pressures
Selection is the only step that reads the fitness function, and it removes diversity. Variation creates it. The behaviour of the whole method is the balance between them.
Scale, Not Just Ranking
Scores of 10 and 9 give the winner 0.53 of the parent slots. Add 990 to both — nothing about which is better changes — and its share falls to 0.5003.
Too Much Pressure
- The best individual takes most of the parent slots
- Its copies fill the population in a few generations
- Crossover of identical parents returns the parent
- Only mutation is left: a slow random walk
- And it looks converged rather than stuck
Where the GA Sits
- Random search — no memory, but a real baseline
- Grid search — systematic, and multiplies per dimension
- Hill climbing — one point, first local optimum wins
- Simulated annealing — one point, sometimes accepts worse
- Genetic algorithm — a population, plus recombination
A Strategy, Not a Model
A genetic algorithm produces no model. It chooses — features, hyperparameters, architectures — and your fitness function does the training. If the objective is differentiable, use the gradient.