The Transformer Architecture
How stacking attention and feed-forward layers — with residuals, layer norm, and positional encoding — produced the architecture that now dominates all of deep learning.
Attention Is All You Need
The paper removed recurrence entirely. No hidden states, no sequential dependencies — only attention layers and feed-forward networks. Within a year, it surpassed every RNN benchmark. Within two, it reshaped the whole field.
Encoder & Decoder Stacks
- Encoder: N layers of self-attention + FFN, processes full input in parallel
- Decoder: N layers of masked self-attention + cross-attention + FFN
- Masked self-attention prevents attending to future tokens (autoregressive)
- Cross-attention lets decoder query encoder representations at each step
Injecting Order
Added to token embeddings before the first layer. Sinusoidal form generalizes to lengths unseen in training. Modern models often use learned or relative encodings (RoPE, ALiBi).
Training Deep Stacks
Residual connections give gradients a direct highway through deep networks. Layer norm stabilizes activation scales independently of batch size.
Per-Position Computation
Attention mixes across positions. FFN transforms each position independently. Expands to 4×dmodel then contracts. Research shows FFN layers store factual knowledge.
Full Parallelism
- RNNs: sequential dependency forces ht to wait for ht-1
- Transformers: all positions computed simultaneously as matrix multiplies
- Training throughput scales with GPU parallelism, not sequence length
- More hardware = proportionally more throughput
Long-Range Dependencies
- RNN: gradient travels through k steps to connect position t to t−k
- Transformer: every position attends to every other in a single layer (path length = 1)
- Long-range dependencies are no harder to learn than short-range ones
- Even LSTM degrades at very long sequences; Transformer does not
Predictable Improvement
- Performance follows a power law in compute, data, and parameters
- Doubling any resource yields a consistent, predictable gain
- This regularity enabled GPT-3, GPT-4, and modern foundation models
- Main bottleneck: O(n²) attention — FlashAttention and sparse methods address this
What You Learned
The Transformer stacks attention + FFN + residuals + layer norm into a fully parallel architecture. Positional encoding injects order. Three properties — parallelism, constant path length, and power-law scaling — explain its dominance. Module 9.3 surveys the modern Transformer family: BERT, GPT, T5, and ViT.