The AI Revolution

From rule-based systems to machines that write poetry — how did we get here?

What is Artificial Intelligence?

AI is the science of making machines do things that would require intelligence if done by a human — learning, reasoning, problem-solving, and creating.

📏
Rule-Based (Classical AI)
Explicit if/then rules written by programmers. Powerful but brittle.
📊
Machine Learning
Systems that learn patterns from data without being explicitly programmed.
🧠
Deep Learning
Neural networks with many layers that learn abstract representations.
✨
Generative AI
Models that create new content — text, images, code, music.
In Plain English

Simple: A recipe book versus a cook who learned by tasting. The rule-based system follows the recipe exactly and falls over the moment an ingredient is missing; the other three all learned from experience instead, each one from more experience than the last.

Technical: Symbolic (rule-based) AI and machine learning are two approaches under the same umbrella. Deep learning is a subset of machine learning — networks with many layers — and generative AI is a use of deep learning that models a distribution well enough to sample new items from it.

The Early Days of AI

The dream of thinking machines is older than you might think.

📜
1950
Alan Turing — “Can Machines Think?”
Proposed the Turing Test in his landmark paper “Computing Machinery and Intelligence”
Read Paper
🏛️
1956
Dartmouth Conference
The term “Artificial Intelligence” is coined. The field is officially born.
Learn More
💬
1966
ELIZA Chatbot
MIT builds a chatbot that mimics a therapist — surprisingly convincing for its era.
Learn More
❄️
1974–93
The AI Winters
Overpromise, underdeliver. Funding dries up twice as hype cycles collapse.
Learn More
♟️
1997
Deep Blue Beats Kasparov
IBM’s chess computer defeats the reigning world champion in a historic match.
Learn More
In Plain English

Simple: Nearly fifty years pass on this timeline and nothing on it can hold a conversation. The field kept announcing it was nearly there, and kept being wrong — twice badly enough that the money stopped.

Technical: The 1950–1997 arc is symbolic AI: hand-encoded rules and search. Turing (1950) proposed an imitation-game criterion; ELIZA (1966) was pattern substitution with no model of meaning; Deep Blue (1997) was alpha–beta search over a hand-tuned evaluation function. Every win came from search and rules, which is why none of it generalised.

From AI Winter to Spring

For decades, neural networks were a dead end. Three breakthroughs changed everything.

📊
The Data Explosion
~1,000× More Data in a Decade
The internet, smartphones, and sensors created oceans of training data that hungry algorithms needed.
⚡
The GPU Revolution
Orders of Magnitude Faster
Graphics cards designed for games turned out to be perfect for parallel matrix math — the core of neural networks.
🧮
Better Algorithms
ReLU, Dropout & Backprop Fixes
Simple but crucial improvements let networks go deeper without vanishing gradients.
🏆
2012 — The Breakthrough
AlexNet Crushes ImageNet
A deep CNN obliterated the competition by 10+ percentage points. The third AI spring had begun.
Read About AlexNet
In Plain English

Simple: The recipe had been written down for years; what finally arrived was the ingredients and the oven. More photographs than anyone could ever look at, and a chip built for video games that turned out to be perfect for the job.

Technical: Backpropagation dates to the 1980s; what changed by 2012 was scale and conditioning — web-scale labelled corpora such as ImageNet, GPU throughput on dense matrix multiplication, and ReLU plus dropout, which mitigated vanishing gradients and overfitting in deep stacks. AlexNet (Krizhevsky et al., 2012) combined all three.

Machine Learning

Traditional programming: humans write rules. ML flips it — data teaches the rules.

🧪
ML Playground
Feed examples, train a model, test predictions
1
Training Data
Input (x)Output (y)
→
2
Train
Waiting for training data...
3
Predict
f( ) = —
Try the pattern: 2→4, 5→10, 8→16. Can the model discover y = 2x?
Under the hood: The model uses linear regression — it finds the line y = mx + b that best fits your data by minimizing the error between predictions and actual values. Each “training” step adjusts m (slope) and b (intercept) to get closer.
In Plain English

Simple: You are not teaching it arithmetic. You show it 2→4, 5→10, 8→16 and it works out “double it” on its own — the same way you spotted the pattern before anyone told you the rule.

Technical: Least-squares linear regression: the widget fits a slope m and an intercept b by minimising the squared error between predicted and observed y. The learned “rule” is those two numbers — not a stored lookup table of the examples you typed in.

Supervised vs. Unsupervised

The two main approaches to ML. Click the canvases to add points and see the difference.

Supervised
You label the data — the model learns boundaries
Unsupervised
No labels — the model finds clusters on its own
In Plain English

Simple: Two ways to sort a pile of laundry. Either somebody already told you which pile each sock belongs in, or you simply notice that some are woolly and some are not, and make your own piles.

Technical: Supervised learning fits a decision boundary to (x, y) pairs where y is a supplied label. Unsupervised learning has no y and instead partitions the input space by a similarity objective — k-means minimises within-cluster variance. The clusters it finds carry no names, because nothing supplied any.

Neural Networks

Inspired by the brain — layers of “neurons” connected by weighted links.

Input Hidden Output x₁ x₂ x₃ y₁ y₂

The network learns by tuning millions of weights to minimize prediction errors.

In Plain English

Simple: Picture an enormous switchboard where every wire has its own volume knob. Learning is turning millions of those knobs a fraction at a time until the right thing comes out of the far end.

Technical: Each layer computes a weighted sum of its inputs plus a bias, then applies a non-linearity. Training is gradient descent on a loss function: backpropagation supplies the derivative of the error with respect to every single weight, and each weight then takes a small step against its own gradient.

The Perceptron (1958)

Click inputs to toggle, adjust weights — watch the computation flow.

1.0 + − 1.0 + − 0 x₁ 0 x₂ -1.5 bias + − Σ = 0.0 step(x) 0 output
0×1.0 + 0×1.0 + (-1.5) = -1.5 → 0
Click the circles to toggle inputs • Use +/− to adjust weights
In Plain English

Simple: A doorman with one rule: add up how much each thing matters to you, and if the total clears the bar, let it in. Change how much things matter and you change who gets in.

Technical: The perceptron (Rosenblatt, 1958) computes a step function of w·x + b — a single linear threshold unit. Adjusting the weights and the bias moves and rotates one hyperplane, which is exactly why it can realise AND and OR but never XOR: XOR is not linearly separable.

Deep Learning

Stack many layers and deep networks learn increasingly abstract representations. Each layer builds on the one before it.

Layer
Sees
Example (Faces)
1
Raw pixels
Edges, gradients
2
Edge combos
Shapes, curves
3
Object parts
Eyes, noses
4
Whole objects
Faces, identity
~1.8T
GPT-4 parameters — outside estimate; OpenAI has never published this
120+
Layers deep — reported
~13T
Training tokens — reported
Why depth matters: A shallow network can only learn simple patterns. Stacking layers lets the model compose features — from pixels to edges to shapes to concepts. This compositionality is what makes deep learning powerful enough for language, vision, and code generation.
In Plain English

Simple: Nobody teaches a child that an eye is two curved lines and a dot. They see enough faces and the parts assemble themselves — edges first, then shapes, then whole faces.

Technical: Depth buys composition. Layer k builds its features out of layer k−1’s features, so expressive power grows with depth rather than width; a shallow network needs exponentially more units to express the same functions. The edges → parts → objects hierarchy in the table was observed empirically in trained convolutional networks.

More Than Autocomplete

Your phone guesses the next word. An LLM continues the whole thought.

Autocomplete vs. LLM
📱 Phone autocomplete
Pick a phrase above.
Predicts one likely word from local word statistics.
✨ Large Language Model
Pick a phrase above.
Continues a coherent thought across many tokens.
In Plain English

Simple: Your phone finishes the word. The model finishes the thought — and then the paragraph after it, still on the same subject.

Technical: Phone autocomplete is typically a short-order n-gram or a lightweight model over local word statistics. An LLM conditions every token on the entire preceding context and feeds its own output back in, so coherence is maintained across hundreds of tokens rather than one.

The Attention Mechanism

When processing a word, the model “pays attention” to other relevant words — here is the recipe, then the thing itself.

🎯
1. Score every word
Reading one word, the model rates how relevant each other word in the sentence is to it.
📊
2. Turn scores into shares
The ratings are squashed into proportions that add up to 100% — the attention weights.
🪨
3. Blend the meanings
The word’s representation becomes a mixture of the words it weighted most heavily.
Because every word is scored against every other word, a pronoun can reach back to its noun no matter how far away it sits. Try it below: hover “it” and watch the three steps above run on one sentence. The formula behind the weights arrives in Module 2.
Attention Spotlight — what does “it” refer to?
Hover or tap “it” to see which words it attends to.
Self-attention resolves “it” by placing the most weight on “cat” — so the model knows the cat was tired, not the mat. Shown here for a bidirectional encoder, which may look both ways: “it” also draws weight from “was” and “tired”, which come after it. A GPT-style decoder is causal — it can only look left. These weights are illustrative and hand-authored, not read out of a trained model — but the effect is real and published: Vaswani et al. (2017) show it in Figure 4, resolving “its” to “Law” at layer 5 of 6.
In Plain English

Simple: “The cat sat on the mat because it was tired.” A mat cannot be tired. You settled that without noticing you had done it — you looked back at exactly one earlier word and ignored the rest. The model has to earn it, and above you can watch the moment it does.

Technical: For each position the model computes a relevance score against every other position, normalises those scores into a distribution that sums to 1, and takes the weighted sum of the other positions’ representations. Coreference then falls out of those weights rather than a separate resolution stage: the query at “it” scores highest against “cat”. Because every pair is scored directly, the path between any two tokens is one step regardless of distance.

Why Transformers Win

RNNs read one word at a time. Transformers see all words at once.

► RNN — Sequential
—
The
→
cat
→
sat
→
on
→
the
→
mat
✓
★ Transformer — Parallel
—
The
cat
sat
on
the
mat
✓
Press Race! to see the difference
RNN Bottleneck
Each word must wait for the previous one to finish. Word 6 can’t start until words 1–5 are done.
Transformer Advantage
All words are processed simultaneously using self-attention. No waiting, no bottleneck.
Further watching: 3Blue1Brown — a visual intro to transformers (external video)
In Plain English

Simple: Six people passing a message down a line, one at a time, against six people reading the same page at once. The second group finishes in the time the first spends on a single hand-off.

Technical: An RNN’s hidden state at step t depends on step t−1, so training cannot be parallelised along the sequence. Self-attention has no such recurrence: every position is computed in one matrix multiplication, trading O(n) sequential steps for O(n²) pairwise comparisons that a GPU performs simultaneously (Vaswani et al., 2017).

Three Transformer Families

Same transformer, wired three ways for three different jobs — a writer, a reader, and a translator. Toggle to see which blocks each one uses. The original 2017 Transformer was the full shape — six encoder layers and six decoder layers (Vaswani et al., §3.1); the single-stack families came afterwards.

Input tokens Encoder bidirectional stack × N Decoder causal (masked) stack × N Output next token
In Plain English

Simple: The same engine in three different chassis. One is built to keep moving forwards (writing), one to take in the whole page at once (reading), and one to receive at one end and produce at the other (translating).

Technical: Decoder-only (GPT) stacks masked self-attention and is trained to predict the next token. Encoder-only (BERT) uses unmasked self-attention and is trained to fill masked positions. Encoder–decoder (T5) keeps both stacks plus cross-attention, mapping an input sequence to a separately generated output sequence.

From GPT to Claude

The transformer ignited a race that reshaped the tech industry in just 7 years.

📄
2017
“Attention Is All You Need”
Vaswani et al. drop recurrence altogether — 6 encoder and 6 decoder layers, 8 attention heads — and every model below this line is built on it.
arXiv:1706.03762
🔬
2018
GPT-1 — 117M Parameters
OpenAI proves transformers can generate coherent, contextual text from pre-training alone.
Read Paper
🚀
2020
GPT-3 — 175B Parameters
Few-shot learning emerges: give examples in the prompt, the model generalizes. A paradigm shift.
Read Paper
💥
Nov 2022
ChatGPT Goes Viral
100M users in 2 months — the fastest-growing consumer product in history.
OpenAI
⚔️
2023
Claude, Gemini & Llama
Anthropic, Google, and Meta enter the race. Open and closed-source frontier models compete.
🤖
2025
Agentic AI
Models that take action autonomously — writing code, running tests, deploying to production.

How big are the models today?

Two different worlds, and only one of them publishes its numbers.

Tier
Size
Where it runs, and who ships it
Tiny / edge
1B–4B
On the device itself — a phone or a laptop, no network. Small Gemma, Phi and Llama variants.
Small / medium
7B–70B
The workhorses: strong enough for real work, cheap enough to self-host. Most Llama, Qwen and Mistral releases.
Open flagship
120B–685B
Server-class open weights that approach closed-model reasoning. Qwen and DeepSeek variants; weights you can download.
Closed frontier
~1T–2T+ (estimated)
API only. OpenAI, Google and Anthropic do not publish parameter counts — every number here is an outside estimate.
Mixture-of-Experts (MoE), and why the big number is misleading. A frontier model is usually not one dense network. It is many sub-networks — experts — with a small router that picks a couple of them per token. So a model advertised at 685B parameters may only run ~37B of them for any given token. That is why an open flagship can rival a closed one without the running cost the headline suggests: total parameters is a size, active parameters is the bill. Always ask which one is being quoted.
In Plain English

Simple: The same idea, made bigger every year, until at some point it stopped being a research demo and became something a hundred million people were using before anyone had finished arguing about it.

Technical: The arc is scale plus interface. GPT-1 (2018) showed that generative pre-training transfers; GPT-3 (2020) showed in-context few-shot learning emerging at around 175B parameters; ChatGPT (2022) added instruction tuning and a chat surface — a product change, not an architectural one. Time-sensitive: the 2023 and 2025 rows describe a landscape that moves faster than any slide.

Knowledge Check

Seven questions on this module. Answer to see why — the explanation appears whether you were right or wrong.

Question 1 of 0
Score 0/0

Key Takeaways

📜
AI Is Not New
The dream started in 1950 — cheap compute and big data made it practical.
🔄
ML Flips Programming
Data + Answers = Rules. Let the machine discover patterns.
🏷️
Two Learning Types
Supervised (labeled) vs. Unsupervised (pattern discovery).
⚡
Transformers Won
Attention + parallel reading = nearly every modern language model. Writing is still one token at a time.
Machine LearningNeural NetworksSupervised LearningTransformersAttentionDeep Learning