Modern Transformer Models
One architecture, many forms — how the Transformer blueprint became BERT, GPT, T5, and ViT, reshaping NLP and computer vision alike.
BERT — Encoder Only
Bidirectional Encoder Representations from Transformers. Pretrained with masked language modeling: predict randomly hidden tokens from full left + right context. Learns deep bidirectional representations ideal for understanding tasks.
Pretrain Once, Fine-Tune Many
- RoBERTa: more data, longer training, no NSP — significantly stronger
- DistilBERT: 40% smaller, 60% faster, 97% of BERT's performance
- ALBERT: parameter sharing across layers, much smaller footprint
- DeBERTa: disentangled attention (content vs. position) — GLUE/SuperGLUE SOTA
GPT — Decoder Only
Trained autoregressively to predict the next token. GPT-2 showed zero-shot task performance at 1.5B params. GPT-3 at 175B parameters revealed in-context learning: few examples in the prompt teach complex tasks without gradient updates.
RLHF & Instruction Tuning
- Raw LMs predict text well but ignore intent
- RLHF: human raters rank outputs; PPO optimizes for preference
- InstructGPT → ChatGPT → GPT-4: same idea at growing scale
- Instruction tuning (no RL) achieves most gains more simply
T5 — Text-to-Text
Full encoder–decoder Transformer. Every task becomes text-in, text-out: translation, classification, QA, summarization — same model, same loss, same fine-tuning. Pretrained on C4 with span corruption. Flan-T5 adds 1,800+ instruction tasks.
ViT — Vision Transformer
Divide image into 16×16 patches, embed each, add positional encoding, run a standard Transformer encoder. No convolutions. No local bias. At small scale CNNs win; at large scale ViT surpasses them — the same scale story as GPT-3.
Swin, DeiT, DINO
- DeiT: trains ViT on ImageNet alone with distillation — no giant dataset needed
- Swin: hierarchical patch merging + local windows → multi-scale features for detection/segmentation
- DINO/DINOv2: self-supervised ViT features generalize to depth, segmentation, retrieval with no labels
Pretrain & Adapt
- Large Transformers trained on broad data become universal starting points
- CLIP: vision + language contrastive learning → zero-shot image classification
- GPT-4V, LLaVA: language models extended to image inputs
- Adaptation strategies: fine-tuning, LoRA, prompt engineering, RAG
What You Learned
BERT (encoder-only, MLM), GPT (decoder-only, autoregressive), T5 (encoder–decoder, text-to-text), and ViT (patches + Transformer encoder) are all one architecture, differently configured. Foundation models and adapt-don’t-retrain is now the standard paradigm.