ML 101
M09 · L02
Module 9

The Transformer Architecture

How stacking attention and feed-forward layers — with residuals, layer norm, and positional encoding — produced the architecture that now dominates all of deep learning.

01 / 11
ML 101
M09 · L02
Vaswani et al., 2017

Attention Is All You Need

The paper removed recurrence entirely. No hidden states, no sequential dependencies — only attention layers and feed-forward networks. Within a year, it surpassed every RNN benchmark. Within two, it reshaped the whole field.

Encoder layers
N = 6
Model dim
d = 512
02 / 11
ML 101
M09 · L02
Architecture Overview

Encoder & Decoder Stacks

  • Encoder: N layers of self-attention + FFN, processes full input in parallel
  • Decoder: N layers of masked self-attention + cross-attention + FFN
  • Masked self-attention prevents attending to future tokens (autoregressive)
  • Cross-attention lets decoder query encoder representations at each step
03 / 11
ML 101
M09 · L02
Positional Encoding

Injecting Order

Sinusoidal PE
PE_{(pos,2i)}=\sin\!\left(\frac{pos}{10000^{2i/d}}\right)

Added to token embeddings before the first layer. Sinusoidal form generalizes to lengths unseen in training. Modern models often use learned or relative encodings (RoPE, ALiBi).

04 / 11
ML 101
M09 · L02
Residual + LayerNorm

Training Deep Stacks

Sublayer Wrapper
\text{LayerNorm}(x+\text{Sublayer}(x))

Residual connections give gradients a direct highway through deep networks. Layer norm stabilizes activation scales independently of batch size.

05 / 11
ML 101
M09 · L02
Feed-Forward Sublayer

Per-Position Computation

FFN
\max(0,xW_1+b_1)\,W_2+b_2

Attention mixes across positions. FFN transforms each position independently. Expands to 4×dmodel then contracts. Research shows FFN layers store factual knowledge.

06 / 11
ML 101
M09 · L02
Why Transformers Scale

Full Parallelism

  • RNNs: sequential dependency forces ht to wait for ht-1
  • Transformers: all positions computed simultaneously as matrix multiplies
  • Training throughput scales with GPU parallelism, not sequence length
  • More hardware = proportionally more throughput
07 / 11
ML 101
M09 · L02
Constant Path Length

Long-Range Dependencies

  • RNN: gradient travels through k steps to connect position t to t−k
  • Transformer: every position attends to every other in a single layer (path length = 1)
  • Long-range dependencies are no harder to learn than short-range ones
  • Even LSTM degrades at very long sequences; Transformer does not
08 / 11
ML 101
M09 · L02
Scaling Laws

Predictable Improvement

  • Performance follows a power law in compute, data, and parameters
  • Doubling any resource yields a consistent, predictable gain
  • This regularity enabled GPT-3, GPT-4, and modern foundation models
  • Main bottleneck: O(n²) attention — FlashAttention and sparse methods address this
09 / 11
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 11
ML 101
Summary
Recap

What You Learned

The Transformer stacks attention + FFN + residuals + layer norm into a fully parallel architecture. Positional encoding injects order. Three properties — parallelism, constant path length, and power-law scaling — explain its dominance. Module 9.3 surveys the modern Transformer family: BERT, GPT, T5, and ViT.

11 / 11