Reading
Stories Mode

The Transformer Architecture

~26 min read Lesson 2 of 3 in Module 9

Attention Is All You Need

In 2017, Vaswani et al. published “Attention Is All You Need,” a paper whose title was a provocation as much as a description. The authors argued that the recurrent connections that had defined sequence modeling for three decades — the hidden states threading information forward one step at a time — were not just unnecessary but actively harmful. All the work they did could be done better, faster, and more scalably with attention alone.

The result was the Transformer: a model built entirely from attention layers and feed-forward networks, with no recurrence whatsoever. Within a year, it had surpassed state-of-the-art RNN and CNN models on every major language benchmark. Within two years, scaled-up Transformer variants (BERT, GPT-2) had reshaped the entire field. The Transformer is now the dominant architecture not only for language but for vision, audio, protein structure prediction, and virtually every domain where sequence or spatial data arises.

This lesson dissects the Transformer architecture from the ground up: the encoder-decoder structure, positional encoding, layer normalization with residual connections, feed-forward sublayers, and the crucial question of why Transformers scale better than RNNs. By the end, you will understand not just how the Transformer works but why each design choice was made.

The Encoder–Decoder Structure

The original Transformer was designed for sequence-to-sequence tasks (specifically machine translation), so it uses an encoder–decoder structure. The encoder reads the entire input sequence and produces a sequence of continuous representations. The decoder generates the output sequence one token at a time, attending to both the encoder representations and the previously generated output tokens.

Each encoder consists of a stack of N identical layers (N = 6 in the original paper). Each layer has two sublayers: a multi-head self-attention sublayer and a position-wise feed-forward sublayer. Every sublayer is wrapped in a residual connection followed by layer normalization. Crucially, every encoder layer processes the full sequence simultaneously — all positions in parallel, without any sequential dependency.

The decoder is similarly structured with N identical layers, but each decoder layer has three sublayers: a masked multi-head self-attention sublayer (masked to prevent attending to future positions), a cross-attention sublayer (attending to the encoder output), and a position-wise feed-forward sublayer. The masking in the first decoder sublayer implements the autoregressive property: during training, each position in the decoder can only attend to earlier positions in the output sequence, since future tokens are not yet known.

Encoder-Only and Decoder-Only Variants

While the original Transformer uses an encoder–decoder structure, modern applications often use only one half. Encoder-only models (e.g., BERT) process the full input bidirectionally and excel at understanding tasks: classification, named entity recognition, question answering. Decoder-only models (e.g., GPT) generate text autoregressively and excel at generation tasks. The encoder–decoder structure remains dominant for translation, summarization, and other tasks that map one sequence to another.

Positional Encoding

Self-attention has no built-in notion of order: the attention formula treats all positions identically regardless of where they appear in the sequence. For language, position matters enormously — “dog bites man” and “man bites dog” have opposite meanings despite using identical tokens. The Transformer solves this by injecting positional information directly into the input embeddings before the first attention layer.

The original paper uses sinusoidal positional encodings: for each position pos and each dimension i of the model, the encoding is a sine or cosine of a frequency that decreases geometrically with dimension. Even dimensions use sine; odd dimensions use cosine:

Sinusoidal Positional Encoding
\begin{aligned}PE_{(pos,2i)} &= \sin\!\left(\frac{pos}{10000^{2i/d_{\text{model}}}}\right)\\ PE_{(pos,2i+1)} &= \cos\!\left(\frac{pos}{10000^{2i/d_{\text{model}}}}\right)\end{aligned}
Each position pos gets a unique vector of dmodel values. Different dimensions encode position at different frequencies — low dimensions capture fine-grained local position, high dimensions capture broad global position. The sinusoidal form allows the model to generalize to sequence lengths not seen during training: the relative position between any two tokens can be expressed as a linear function of their individual encodings.

The positional encoding is added (not concatenated) to the token embedding. The resulting sum passes into the first encoder or decoder layer. Because both the embedding and the positional encoding have the same dimensionality dmodel, the model can learn to disentangle token identity from positional information through the learned weight matrices.

Modern models often use learned positional embeddings instead — a lookup table where each position has a trainable embedding vector. Learned embeddings tend to work slightly better in practice, at the cost of not generalizing beyond the maximum sequence length seen during training. Relative positional encodings (e.g., RoPE, ALiBi) go further, directly modifying attention scores based on the distance between positions rather than adding absolute position signals to embeddings. These relative methods have become standard in large language models.

Residual Connections and Layer Normalization

Two components appear consistently throughout the Transformer and are essential for training deep networks: residual connections and layer normalization. Each sublayer (attention or feed-forward) is wrapped in the pattern:

Sublayer Wrapper
\text{Output} = \text{LayerNorm}\bigl(x + \text{Sublayer}(x)\bigr)
x is the input to the sublayer and Sublayer(x) is the output of the attention or feed-forward operation. The residual connection adds the original input directly to the sublayer output. LayerNorm then normalizes the sum. This order (residual then norm) is the “post-norm” convention of the original paper; modern models often use “pre-norm” (LayerNorm applied before the sublayer), which provides more stable training.

Residual connections address the vanishing gradient problem in deep networks. Without them, gradients must flow through every layer during backpropagation — each multiplicative operation potentially shrinking the gradient toward zero. The residual connection provides a direct path for gradients to flow from the output all the way back to the input without passing through the sublayer computation. This is why very deep networks (100+ layers) can be trained stably with residuals but not without them.

Layer normalization normalizes the activations across the feature dimension (the dmodel dimension) for each position independently. Unlike batch normalization, which normalizes across the batch dimension, layer normalization operates on a single example at a time and is therefore unaffected by batch size. This is critical for sequence models where batch sizes are small and sequences have variable lengths. Layer norm also stabilizes training by keeping the scale of activations consistent across layers.

The Feed-Forward Sublayer

Each Transformer layer also includes a position-wise feed-forward network (FFN): a two-layer MLP applied independently to each position. The FFN expands the representation to a larger intermediate dimension dff (typically 4×dmodel), applies a nonlinearity, then projects back to dmodel:

Position-Wise Feed-Forward Network
\text{FFN}(x) = \max(0,\, xW_1 + b_1)\,W_2 + b_2
The same W1, W2, b1, b2 are applied at every position independently, but they differ across layers. The original paper uses ReLU as the activation; modern models often use GeLU or SwiGLU. With dmodel=512 and dff=2048, the FFN has roughly 4× more parameters than the attention sublayer, making it the largest parameter block in each layer.

The FFN sublayer serves a different function from the attention sublayer. Attention aggregates information across positions — it is fundamentally a mixing operation. The FFN processes each position independently — it is a per-position nonlinear transformation. Together, they form a complementary pair: attention for cross-position communication, FFN for per-position computation.

Research into large language models has found that the FFN layers store and retrieve factual knowledge. Individual neurons in the FFN can be identified that activate for specific facts (e.g., “Paris is the capital of France”), and editing these neurons changes the model’s factual claims. This “knowledge neurons” view suggests the FFN acts as a form of associative memory layered on top of the attention mechanism’s structural reasoning.

Why Transformers Scale Better Than RNNs

The Transformer’s dominance over RNNs is not just about performance on a fixed dataset — it is about scaling behavior. As you increase the amount of data, compute, and model parameters, Transformers continue to improve in ways that RNNs do not. Three structural properties explain this:

1. Parallelism during training. RNNs process sequences step by step: computing ht requires ht−1, which requires ht−2, and so on. The sequential dependency means you cannot parallelize across the time dimension — you must wait for each step to complete before starting the next. With a sequence of length 1000, this is 1000 sequential operations regardless of how many GPU cores you have. Transformers compute all positions simultaneously as matrix multiplications. Training throughput scales with hardware parallelism rather than being bottlenecked by sequence length.

2. Constant path length between any two positions. In an RNN, the gradient must travel through every intermediate hidden state to connect position t to position t−k — a path length of k. For long sequences, this path is so long that gradients vanish and the model fails to learn long-range dependencies (even LSTM ameliorates but does not eliminate this). In a Transformer, every position attends directly to every other position in a single layer — path length 1. Long-range dependencies are no harder to learn than short-range ones.

3. Predictable scaling laws. Kaplan et al. (2020) found that Transformer language model performance follows a power law in compute, dataset size, and parameter count: doubling any one resource yields a predictable, consistent improvement. This regularity means you can estimate in advance how much performance you will gain from scaling up. RNNs do not exhibit such clean scaling behavior, in part because their sequential nature caps the benefit of additional compute at a given sequence length. The Transformer’s architecture is uniquely amenable to the “scale everything” strategy that produced GPT-3, GPT-4, and similar models.

The O(n²) Bottleneck

The Transformer’s main limitation is the quadratic cost of full self-attention — the same O(n²) cost introduced when self-attention was defined in the previous lesson: computing the n×n attention matrix requires O(n²d) time and O(n²) memory. For moderate sequence lengths (512–2048), this is manageable. For long documents, high-resolution images, or genomic sequences (tens of thousands of tokens), the cost becomes prohibitive. Research into sparse attention (BigBird, Longformer), linear attention (Performer, RWKV), and hardware-efficient algorithms (FlashAttention) addresses this limitation while preserving most of the Transformer’s expressiveness. FlashAttention in particular achieves the exact same output as full attention with O(n) memory by fusing the attention computation into a single kernel, and has become standard in modern LLM training.

The Complete Forward Pass

Let us trace how a sequence flows through the full Transformer encoder. Input tokens are converted to embeddings of dimension dmodel. Positional encodings are added element-wise. The resulting matrix (shape: seq_len × dmodel) passes through N encoder layers. Each layer applies: (1) multi-head self-attention with residual and layer norm, then (2) feed-forward network with residual and layer norm. The output of the final encoder layer is a context-enriched representation of every input token, where each token’s embedding now incorporates information gathered from all other tokens across all attention heads and all layers.

The decoder follows a similar path, with the additional cross-attention sublayer in each layer consuming the encoder output. At each decoding step, the decoder produces a probability distribution over the vocabulary, from which the next token is sampled or greedily selected. That token is appended to the output sequence and the decoder runs again for the next step, continuing until an end-of-sequence token is generated or a maximum length is reached.

During training, teacher forcing is used: the correct output tokens (from the ground-truth target sequence) are fed into the decoder at each step rather than the model’s own predictions. The causal mask ensures the decoder cannot see future tokens, making the training procedure valid (each position only conditions on tokens that would actually be available at inference time). The entire forward pass — both encoder and decoder — is executed in parallel for all positions, making training highly efficient.

Key Takeaways
  • The Transformer replaces recurrence entirely with stacked multi-head self-attention and feed-forward layers. The encoder processes the full input in parallel; the decoder generates output autoregressively using masked self-attention and cross-attention to the encoder.
  • Positional encoding injects order information into token embeddings before the first layer. Sinusoidal encodings allow generalization to unseen lengths; learned embeddings often work better in practice; relative positional encodings (RoPE, ALiBi) have become standard in large models.
  • Every sublayer is wrapped in a residual connection plus layer normalization. Residuals provide gradient highways through deep stacks; layer norm stabilizes activation scales and operates independently of batch size.
  • The feed-forward sublayer is a two-layer MLP applied independently at each position, expanding to 4×dmodel then contracting. It provides per-position nonlinear computation complementing the cross-position mixing of attention. FFN layers appear to store factual knowledge as associative memories.
  • Transformers outscale RNNs because (1) all positions train in parallel, (2) path length between any two positions is constant at 1, and (3) performance follows predictable power laws in compute and data. These properties make the Transformer uniquely suited to the large-scale training regime of modern foundation models.
  • The main limitation is O(n²) attention complexity. FlashAttention, sparse attention, and linear attention variants address this while preserving most expressiveness, enabling Transformer-like architectures to handle much longer sequences.
Previous The Attention Mechanism Overview Next Modern Transformer Models