Reading
Stories Mode

Modern Transformer Models

~28 min read Lesson 3 of 3 in Module 9

From Architecture to Ecosystem

The Transformer architecture described in “Attention Is All You Need” was a blueprint, not a finished product. In the years following its publication, researchers discovered that the same architectural core — stacked attention layers, feed-forward sublayers, residual connections, and layer normalization — could be adapted into a diverse family of models, each optimized for a different setting or task.

Some models kept only the encoder, discarding the decoder to enable bidirectional understanding of text. Others kept only the decoder, scaling autoregressive generation to sizes previously unimaginable. Still others restructured the encoder–decoder relationship to handle arbitrary text-to-text tasks. And beyond language, the same architecture migrated into vision, audio, and multimodal domains, demonstrating a generality that no one fully anticipated in 2017.

This lesson surveys the modern Transformer family: BERT and its encoder-only descendants, the GPT lineage of decoder-only models, T5 and the encoder–decoder paradigm, and Vision Transformers (ViT). Understanding each variant — what it includes, what it discards, and why — gives you a mental map of the landscape you will navigate as a practitioner.

BERT: Bidirectional Encoder Representations

In 2018, Devlin et al. at Google released BERT (Bidirectional Encoder Representations from Transformers). The key insight was simple but powerful: for tasks that require understanding — classification, question answering, named entity recognition — you do not need generation. You need to read the entire input at once and produce a rich contextual representation of every token. An encoder-only model, processing tokens bidirectionally (each token can attend to all others), is ideal for this.

BERT introduced masked language modeling (MLM) as a pretraining objective. During pretraining, 15% of input tokens are randomly replaced with a [MASK] token, and the model must predict the original tokens from context. Because the model sees both left and right context simultaneously, it learns deep bidirectional representations that capture syntactic and semantic relationships far richer than anything a unidirectional model can achieve.

BERT also uses a second pretraining objective called next sentence prediction (NSP): given two sentences, predict whether the second follows the first in the original document. This teaches the model inter-sentence coherence. (Later work found NSP provides limited benefit; RoBERTa and subsequent models dropped it while improving MLM training.)

Fine-Tuning BERT

After pretraining on a large corpus (English Wikipedia + BookCorpus, ~3.3B words), BERT is fine-tuned on each downstream task by adding a task-specific head (e.g., a linear classifier on top of the [CLS] token) and training for a few epochs on labeled data. This two-stage paradigm — pretrain once, fine-tune many times — set the standard for NLP and later for other modalities. BERT-large (340M parameters) achieved state-of-the-art results on 11 NLP benchmarks simultaneously at release, a remarkable demonstration of transfer learning.

The BERT family expanded rapidly. RoBERTa (Facebook, 2019) showed that BERT was undertrained: using more data, larger batches, longer training, and removing NSP improved performance substantially without architectural changes. DistilBERT applied knowledge distillation to produce a model 40% smaller and 60% faster than BERT-base while retaining 97% of its performance. ALBERT used parameter sharing across layers to dramatically reduce parameter counts. DeBERTa introduced disentangled attention (separate representations for content and position) and an enhanced mask decoder, achieving state-of-the-art results on GLUE and SuperGLUE as of 2021.

GPT: The Decoder-Only Lineage

While BERT was optimizing for understanding, OpenAI was pursuing a different question: how well can a decoder-only Transformer — trained to predict the next token autoregressively — learn to generate coherent text? The answer, it turned out, was: extraordinarily well, and better with scale.

GPT-1 (2018) demonstrated that unsupervised pretraining on a large text corpus followed by supervised fine-tuning could outperform task-specific models on multiple NLP tasks. The training objective was simple: maximize the likelihood of each token given all preceding tokens. No masking, no bidirectional context — just left-to-right prediction, the same objective used in language modeling since the n-gram era.

GPT-2 (2019) scaled the model to 1.5 billion parameters and demonstrated something surprising: a language model trained only on next-token prediction, with no task-specific fine-tuning, could perform many downstream tasks reasonably well simply by being prompted in the right way. This was the first clear signal of emergent capabilities: abilities that arise from scale rather than explicit training. GPT-2’s text generation was fluid enough to raise concerns about misuse, leading OpenAI to release it in stages.

GPT-3 (2020) scaled to 175 billion parameters, trained on ~300 billion tokens. The leap in capability was qualitative, not just quantitative. GPT-3 could perform complex reasoning, write code, translate languages, summarize documents, and answer questions — all from a few examples provided in the prompt, without any gradient updates. This few-shot in-context learning ability, where the model learns from examples embedded in the prompt itself, demonstrated a form of meta-learning that emerged purely from scale.

Instruction Tuning and RLHF

Raw language models like GPT-3 are good at predicting text but not necessarily at following instructions. InstructGPT (2022) addressed this via reinforcement learning from human feedback (RLHF): human raters ranked model outputs for quality and helpfulness, and the model was fine-tuned with PPO to maximize predicted human preference. This made GPT-3-class models dramatically more useful and aligned. The same technique underlies ChatGPT, GPT-4, and most modern instruction-following LLMs. Instruction tuning — fine-tuning on datasets of (instruction, response) pairs without RL — provides a simpler alternative that achieves much of the same improvement.

T5: Text-to-Text Transfer Transformer

BERT uses an encoder-only architecture optimized for understanding. GPT uses a decoder-only architecture optimized for generation. In 2019, Google released T5 (Text-to-Text Transfer Transformer), which argued for a unified framework: cast every NLP task as a text-to-text problem, and use a full encoder–decoder Transformer for everything.

In T5’s framing, translation is “translate English to German: [source]” → “[translation]”. Classification is “sentiment: [text]” → “positive” or “negative”. Question answering is “question: [q] context: [c]” → “[answer]”. Summarization is “summarize: [document]” → “[summary]”. The same model, same loss function (cross-entropy on output tokens), and same fine-tuning procedure applies to every task. This unification is elegant and practically powerful: it simplifies the engineering pipeline and allows the model to share representations across tasks.

T5 was pretrained on the Colossal Clean Crawled Corpus (C4), a massive filtered version of Common Crawl, using a span corruption objective: randomly selected spans of tokens are replaced with a single sentinel token, and the model must predict the original spans. This is a generalization of BERT’s masked language modeling to a generative setting — the decoder produces the corrupted spans rather than predicting tokens in-place.

T5-11B (11 billion parameters) achieved state-of-the-art results across a wide range of NLP benchmarks. The T5 family also established the importance of careful data curation: C4 filtering choices (deduplication, quality filtering, removal of offensive content) significantly impact downstream model quality. Flan-T5 (2022) fine-tuned T5 on over 1,800 instruction-formatted tasks, producing a model that generalizes exceptionally well to new tasks from instructions alone.

Vision Transformer (ViT)

The Transformer was designed for sequences of discrete tokens. Images are continuous and two-dimensional — a fundamentally different structure. For three years after the original Transformer paper, convolutional neural networks remained dominant in vision. Then, in 2020, Dosovitskiy et al. showed that a pure Transformer, applied directly to sequences of image patches with minimal modification, could match or exceed CNN performance on image classification when trained at sufficient scale. The result was the Vision Transformer (ViT).

The key idea is patch embedding: divide the image into a grid of non-overlapping patches (e.g., 16×16 pixels each), flatten each patch into a vector, and linearly project it to the model dimension dmodel. Append a learnable [CLS] token at the beginning of the sequence. Add positional embeddings (learned 1D positional embeddings over patch positions). Feed the resulting sequence of tokens through a standard Transformer encoder. The [CLS] token representation at the output is used for classification via a linear head.

ViT has no convolutions, no local inductive biases — unlike CNNs, it does not assume that nearby pixels are more relevant than distant ones. This is both a weakness and a strength. At small data scales, CNNs outperform ViT because convolution’s local bias is a useful prior for images. At large data scales (JFT-300M: 300 million images), ViT matches or surpasses state-of-the-art CNNs. The lesson mirrors GPT-3: at sufficient scale, inductive biases become less important than the model’s ability to learn structure from data.

ViT Variants and DINO

DeiT showed that ViT can be trained effectively on ImageNet alone (without massive pretraining datasets) using strong data augmentation and knowledge distillation from a CNN teacher. Swin Transformer introduced hierarchical patch merging and local window attention, giving ViT the multi-scale feature maps CNNs use — essential for dense prediction tasks like detection and segmentation. DINO and DINOv2 use self-supervised learning on ViT, producing visual features that emerge without labels and generalize remarkably across tasks. DINOv2 features trained on internet images produce near-perfect depth estimation, segmentation, and retrieval with only a linear probe — a signal that large-scale self-supervised ViT pretraining is approaching the quality of supervised ImageNet features at any scale.

A Unified View: Foundation Models

BERT, GPT, T5, and ViT are not separate inventions. They are specializations of a single architectural blueprint — the Transformer — adapted to different modalities and training objectives. This convergence toward a single architecture across modalities has a name: foundation models. A foundation model is a large model trained on broad data at scale, which can be adapted to a wide range of downstream tasks.

The practical consequence is that the field has moved from training task-specific models from scratch to adapting pretrained foundation models. Rather than designing a new architecture for each problem, practitioners now choose an appropriate foundation model, select an adaptation strategy (fine-tuning, prompt engineering, retrieval augmentation, or parameter-efficient fine-tuning via LoRA or adapters), and deploy. The quality of the pretrained representations dramatically raises the ceiling for what is achievable with limited labeled data.

Multimodal models extend this further. CLIP (Radford et al., 2021) trained a vision encoder and a text encoder jointly on 400 million image–caption pairs using contrastive learning, learning a shared embedding space where images and their textual descriptions are nearby. CLIP’s image representations transfer extremely well to downstream tasks and enable zero-shot image classification: given an image, classify it by finding the text description that is closest in CLIP embedding space, with no labeled images required. GPT-4V, LLaVA, and similar models extend GPT-class language models to accept image inputs, enabling visual question answering, image captioning, and document understanding.

The trajectory is clear: Transformers are becoming the universal architecture, the same way deep neural networks became the universal function approximator a decade earlier. The question is no longer “should I use a Transformer?” but “which Transformer variant, pretrained on which data, adapted with which technique?”

Key Takeaways
  • BERT uses encoder-only Transformers with masked language modeling pretraining to learn deep bidirectional representations. It is fine-tuned for understanding tasks (classification, QA, NER). RoBERTa, DistilBERT, ALBERT, and DeBERTa are key descendants.
  • GPT uses decoder-only Transformers trained on autoregressive next-token prediction. Scaling from GPT-1 to GPT-3 revealed emergent in-context learning. RLHF alignment (InstructGPT, ChatGPT) makes raw LMs instruction-following. GPT-4 extended this to multimodal inputs.
  • T5 unifies all NLP tasks as text-to-text problems using a full encoder–decoder Transformer pretrained with span corruption on C4. Flan-T5 adds instruction tuning across 1,800+ tasks.
  • Vision Transformer (ViT) applies a standard Transformer encoder to sequences of image patches. It requires large-scale pretraining to match CNNs, but at sufficient scale surpasses them. Swin Transformer adds hierarchy for dense prediction; DINO uses self-supervised learning for strong visual features.
  • Foundation models — large Transformers pretrained on broad data — are now the standard starting point. Adaptation strategies (fine-tuning, LoRA, prompt engineering, retrieval) determine deployment quality. CLIP and multimodal models extend this paradigm across modalities.
  • The Transformer has become the universal architecture. Understanding the encoder-only, decoder-only, and encoder–decoder variants, and when each is appropriate, is now essential knowledge for any ML practitioner.
Previous The Transformer Architecture Overview Next Feature Engineering