ML 101
M09 · L03
Module 9

Modern Transformer Models

One architecture, many forms — how the Transformer blueprint became BERT, GPT, T5, and ViT, reshaping NLP and computer vision alike.

01 / 11
ML 101
M09 · L03
Google, 2018

BERT — Encoder Only

Bidirectional Encoder Representations from Transformers. Pretrained with masked language modeling: predict randomly hidden tokens from full left + right context. Learns deep bidirectional representations ideal for understanding tasks.

Params (large)
340M
Benchmarks
11 SOTA
02 / 11
ML 101
M09 · L03
The BERT Family

Pretrain Once, Fine-Tune Many

  • RoBERTa: more data, longer training, no NSP — significantly stronger
  • DistilBERT: 40% smaller, 60% faster, 97% of BERT's performance
  • ALBERT: parameter sharing across layers, much smaller footprint
  • DeBERTa: disentangled attention (content vs. position) — GLUE/SuperGLUE SOTA
03 / 11
ML 101
M09 · L03
OpenAI, 2018–2020

GPT — Decoder Only

Trained autoregressively to predict the next token. GPT-2 showed zero-shot task performance at 1.5B params. GPT-3 at 175B parameters revealed in-context learning: few examples in the prompt teach complex tasks without gradient updates.

04 / 11
ML 101
M09 · L03
Alignment

RLHF & Instruction Tuning

  • Raw LMs predict text well but ignore intent
  • RLHF: human raters rank outputs; PPO optimizes for preference
  • InstructGPT → ChatGPT → GPT-4: same idea at growing scale
  • Instruction tuning (no RL) achieves most gains more simply
05 / 11
ML 101
M09 · L03
Google, 2019

T5 — Text-to-Text

Full encoder–decoder Transformer. Every task becomes text-in, text-out: translation, classification, QA, summarization — same model, same loss, same fine-tuning. Pretrained on C4 with span corruption. Flan-T5 adds 1,800+ instruction tasks.

06 / 11
ML 101
M09 · L03
Google Brain, 2020

ViT — Vision Transformer

Divide image into 16×16 patches, embed each, add positional encoding, run a standard Transformer encoder. No convolutions. No local bias. At small scale CNNs win; at large scale ViT surpasses them — the same scale story as GPT-3.

Pretraining data
JFT-300M: 300 million images
07 / 11
ML 101
M09 · L03
ViT Family

Swin, DeiT, DINO

  • DeiT: trains ViT on ImageNet alone with distillation — no giant dataset needed
  • Swin: hierarchical patch merging + local windows → multi-scale features for detection/segmentation
  • DINO/DINOv2: self-supervised ViT features generalize to depth, segmentation, retrieval with no labels
08 / 11
ML 101
M09 · L03
Foundation Models

Pretrain & Adapt

  • Large Transformers trained on broad data become universal starting points
  • CLIP: vision + language contrastive learning → zero-shot image classification
  • GPT-4V, LLaVA: language models extended to image inputs
  • Adaptation strategies: fine-tuning, LoRA, prompt engineering, RAG
09 / 11
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 11
ML 101
Summary
Recap

What You Learned

BERT (encoder-only, MLM), GPT (decoder-only, autoregressive), T5 (encoder–decoder, text-to-text), and ViT (patches + Transformer encoder) are all one architecture, differently configured. Foundation models and adapt-don’t-retrain is now the standard paradigm.

11 / 11