The AI Revolution
From rule-based systems to machines that write poetry — how did we get here?
What is Artificial Intelligence?
AI is the science of making machines do things that would require intelligence if done by a human — learning, reasoning, problem-solving, and creating.
Simple: A recipe book versus a cook who learned by tasting. The rule-based system follows the recipe exactly and falls over the moment an ingredient is missing; the other three all learned from experience instead, each one from more experience than the last.
Technical: Symbolic (rule-based) AI and machine learning are two approaches under the same umbrella. Deep learning is a subset of machine learning — networks with many layers — and generative AI is a use of deep learning that models a distribution well enough to sample new items from it.
The Early Days of AI
The dream of thinking machines is older than you might think.
Simple: Nearly fifty years pass on this timeline and nothing on it can hold a conversation. The field kept announcing it was nearly there, and kept being wrong — twice badly enough that the money stopped.
Technical: The 1950–1997 arc is symbolic AI: hand-encoded rules and search. Turing (1950) proposed an imitation-game criterion; ELIZA (1966) was pattern substitution with no model of meaning; Deep Blue (1997) was alpha–beta search over a hand-tuned evaluation function. Every win came from search and rules, which is why none of it generalised.
From AI Winter to Spring
For decades, neural networks were a dead end. Three breakthroughs changed everything.
Simple: The recipe had been written down for years; what finally arrived was the ingredients and the oven. More photographs than anyone could ever look at, and a chip built for video games that turned out to be perfect for the job.
Technical: Backpropagation dates to the 1980s; what changed by 2012 was scale and conditioning — web-scale labelled corpora such as ImageNet, GPU throughput on dense matrix multiplication, and ReLU plus dropout, which mitigated vanishing gradients and overfitting in deep stacks. AlexNet (Krizhevsky et al., 2012) combined all three.
Machine Learning
Traditional programming: humans write rules. ML flips it — data teaches the rules.
y = mx + b that best fits your data by minimizing the error between predictions and actual values. Each “training” step adjusts m (slope) and b (intercept) to get closer.
Simple: You are not teaching it arithmetic. You show it 2→4, 5→10, 8→16 and it works out “double it” on its own — the same way you spotted the pattern before anyone told you the rule.
Technical: Least-squares linear regression: the widget fits a slope m and an intercept b by minimising the squared error between predicted and observed y. The learned “rule” is those two numbers — not a stored lookup table of the examples you typed in.
Supervised vs. Unsupervised
The two main approaches to ML. Click the canvases to add points and see the difference.
Simple: Two ways to sort a pile of laundry. Either somebody already told you which pile each sock belongs in, or you simply notice that some are woolly and some are not, and make your own piles.
Technical: Supervised learning fits a decision boundary to (x, y) pairs where y is a supplied label. Unsupervised learning has no y and instead partitions the input space by a similarity objective — k-means minimises within-cluster variance. The clusters it finds carry no names, because nothing supplied any.
Neural Networks
Inspired by the brain — layers of “neurons” connected by weighted links.
The network learns by tuning millions of weights to minimize prediction errors.
Simple: Picture an enormous switchboard where every wire has its own volume knob. Learning is turning millions of those knobs a fraction at a time until the right thing comes out of the far end.
Technical: Each layer computes a weighted sum of its inputs plus a bias, then applies a non-linearity. Training is gradient descent on a loss function: backpropagation supplies the derivative of the error with respect to every single weight, and each weight then takes a small step against its own gradient.
The Perceptron (1958)
Click inputs to toggle, adjust weights — watch the computation flow.
Simple: A doorman with one rule: add up how much each thing matters to you, and if the total clears the bar, let it in. Change how much things matter and you change who gets in.
Technical: The perceptron (Rosenblatt, 1958) computes a step function of w·x + b — a single linear threshold unit. Adjusting the weights and the bias moves and rotates one hyperplane, which is exactly why it can realise AND and OR but never XOR: XOR is not linearly separable.
Deep Learning
Stack many layers and deep networks learn increasingly abstract representations. Each layer builds on the one before it.
Simple: Nobody teaches a child that an eye is two curved lines and a dot. They see enough faces and the parts assemble themselves — edges first, then shapes, then whole faces.
Technical: Depth buys composition. Layer k builds its features out of layer k−1’s features, so expressive power grows with depth rather than width; a shallow network needs exponentially more units to express the same functions. The edges → parts → objects hierarchy in the table was observed empirically in trained convolutional networks.
More Than Autocomplete
Your phone guesses the next word. An LLM continues the whole thought.
Simple: Your phone finishes the word. The model finishes the thought — and then the paragraph after it, still on the same subject.
Technical: Phone autocomplete is typically a short-order n-gram or a lightweight model over local word statistics. An LLM conditions every token on the entire preceding context and feeds its own output back in, so coherence is maintained across hundreds of tokens rather than one.
The Attention Mechanism
When processing a word, the model “pays attention” to other relevant words — here is the recipe, then the thing itself.
Simple: “The cat sat on the mat because it was tired.” A mat cannot be tired. You settled that without noticing you had done it — you looked back at exactly one earlier word and ignored the rest. The model has to earn it, and above you can watch the moment it does.
Technical: For each position the model computes a relevance score against every other position, normalises those scores into a distribution that sums to 1, and takes the weighted sum of the other positions’ representations. Coreference then falls out of those weights rather than a separate resolution stage: the query at “it” scores highest against “cat”. Because every pair is scored directly, the path between any two tokens is one step regardless of distance.
Why Transformers Win
RNNs read one word at a time. Transformers see all words at once.
Simple: Six people passing a message down a line, one at a time, against six people reading the same page at once. The second group finishes in the time the first spends on a single hand-off.
Technical: An RNN’s hidden state at step t depends on step t−1, so training cannot be parallelised along the sequence. Self-attention has no such recurrence: every position is computed in one matrix multiplication, trading O(n) sequential steps for O(n²) pairwise comparisons that a GPU performs simultaneously (Vaswani et al., 2017).
Three Transformer Families
Same transformer, wired three ways for three different jobs — a writer, a reader, and a translator. Toggle to see which blocks each one uses. The original 2017 Transformer was the full shape — six encoder layers and six decoder layers (Vaswani et al., §3.1); the single-stack families came afterwards.
Simple: The same engine in three different chassis. One is built to keep moving forwards (writing), one to take in the whole page at once (reading), and one to receive at one end and produce at the other (translating).
Technical: Decoder-only (GPT) stacks masked self-attention and is trained to predict the next token. Encoder-only (BERT) uses unmasked self-attention and is trained to fill masked positions. Encoder–decoder (T5) keeps both stacks plus cross-attention, mapping an input sequence to a separately generated output sequence.
From GPT to Claude
The transformer ignited a race that reshaped the tech industry in just 7 years.
How big are the models today?
Two different worlds, and only one of them publishes its numbers.
Simple: The same idea, made bigger every year, until at some point it stopped being a research demo and became something a hundred million people were using before anyone had finished arguing about it.
Technical: The arc is scale plus interface. GPT-1 (2018) showed that generative pre-training transfers; GPT-3 (2020) showed in-context few-shot learning emerging at around 175B parameters; ChatGPT (2022) added instruction tuning and a chat surface — a product change, not an architectural one. Time-sensitive: the 2023 and 2025 rows describe a landscape that moves faster than any slide.
Knowledge Check
Seven questions on this module. Answer to see why — the explanation appears whether you were right or wrong.