From Fixed Vectors to Dynamic Lookup
At the end of Module 8 we saw how the encoder–decoder architecture compresses an entire input sequence into a single fixed-size context vector, then uses that vector to initialize the decoder. This design is elegant but carries a fundamental bottleneck: for a long input sequence, every nuance of meaning must be squeezed into a vector of perhaps 512 numbers. Critical information is inevitably lost as sequence length grows, and translation quality degrades accordingly.
The attention mechanism, introduced by Bahdanau, Cho, and Bengio in 2015, solved this problem by discarding the single-vector bottleneck entirely. Instead of compressing the input once and for all, the decoder is given direct access to every encoder hidden state at every decoding step. It selectively focuses on whichever encoder states are most relevant to generating the current output token — paying attention to some positions more than others. This soft, differentiable lookup turned out to be one of the most consequential ideas in modern machine learning.
This lesson develops the attention mechanism from first principles: the intuition of weighted retrieval, the formal Query–Key–Value abstraction, scaled dot-product attention, multi-head attention, self-attention, and cross-attention. These concepts form the complete vocabulary needed to understand the Transformer architecture in the next lesson.
Attention as Weighted Information Retrieval
Think of a library with N books (the encoder hidden states h1, …, hN). You arrive with a question (the decoder’s current state st) and want to retrieve the most relevant information. A naive approach would read all books with equal weight. A better approach would assign each book a relevance score, normalize those scores into a probability distribution, and form a weighted summary of the content — reading the most relevant books carefully while skimming the rest.
This is precisely what attention does. The alignment score etj measures how well the decoder’s current state matches encoder state hj. A softmax over all alignment scores produces attention weights αtj that sum to 1 across all encoder positions j. The context vector ct is the weighted sum of encoder states under those weights. The decoder then generates the next output token using both its own hidden state st and this context vector ct.
Crucially, the attention weights are differentiable: gradients flow back through the softmax and the alignment network, training the model to produce good attention patterns without any explicit supervision on where to look. The model learns which positions to attend to purely from the task objective (e.g., next-token prediction or translation accuracy).
Bahdanau-style attention is soft attention: the context vector is a weighted average over all positions, so every encoder state contributes (albeit with different weights). Hard attention selects exactly one encoder state at each step — a discrete, non-differentiable operation. While hard attention is more interpretable, it requires reinforcement learning or the REINFORCE trick to train, which is less stable. Soft attention became the dominant approach because it is fully differentiable and integrates seamlessly with backpropagation.
The Query–Key–Value Abstraction
Bahdanau’s original formulation used a learned alignment network to compute relevance scores. Vaswani et al. (2017) — the “Attention Is All You Need” paper — recast attention in a more general and computationally efficient form using three linear projections: queries (Q), keys (K), and values (V).
The intuition comes from information retrieval: when you search a database, your query is compared against keys in the database, and matching keys return their associated values. In neural attention, all three are continuous vectors learned by the model. Each input position produces a key (what I contain) and a value (what I return if selected). Each query position asks (what am I looking for?). Relevance is measured as the dot product between a query and each key. The output is the weighted sum of values, where weights come from the query–key similarities.
Concretely, given input matrices X (shape: sequence length × model dimension dmodel), three learned weight matrices WQ, WK, WV project the inputs into query, key, and value spaces: Q = XWQ, K = XWK, V = XWV. This linear projection is computationally inexpensive and enables all positions to be processed in parallel as matrix operations — a fundamental advantage over sequential RNNs.
Scaled Dot-Product Attention
With Q, K, V in hand, the attention output is computed as:
The scaling factor √dk is critical. When dk is large (e.g., 64 or 128), random dot products between unit-normal query and key vectors have variance dk — their magnitude grows with dimension. Without scaling, these large values make the softmax extremely peaked (near a one-hot), collapsing the gradient and making training slow. Dividing by √dk restores unit variance and keeps gradients healthy throughout training.
An optional mask can be added before the softmax. For causal (autoregressive) attention, the mask sets future positions to −∞ so that each position can only attend to earlier positions. This implements the autoregressive property required for language generation: the model cannot “see” future tokens when predicting the current one. For padding, a mask sets padded positions to −∞ so they do not contribute to the attention weights.
One reason scaled dot-product attention became dominant is computational efficiency. Computing QK³ is a single matrix multiply — a highly optimized operation on modern GPUs/TPUs. The alignment network in Bahdanau attention requires evaluating a feed-forward network for each pair of (decoder position, encoder position), making it O(Tout × Tin) sequential operations. Scaled dot-product attention computes all similarities in parallel as a matrix multiply, enabling highly parallelized training across all sequence positions simultaneously.
Multi-Head Attention
A single attention head can learn one kind of relevance relationship — for example, syntactic agreement or semantic similarity. But language and other sequential data are structured in many overlapping ways simultaneously. Multi-head attention runs h independent attention operations in parallel, each with its own Q, K, V projections, allowing the model to jointly attend to information from different representation subspaces.
Empirically, different attention heads specialize in different linguistic patterns. In trained language models, some heads track syntactic structure (subject–verb agreement, coreference), others track positional proximity, and others attend to semantically related words regardless of position. This emergent specialization arises purely from training on next-token prediction — no explicit supervision is provided about what each head should do.
The concatenation of h head outputs (each of dimension dk) produces a vector of dimension h×dk = dmodel. The final projection WO mixes information across heads, allowing the model to combine the different relationship patterns each head discovered into a unified representation for the next layer.
Self-Attention: Attending to Your Own Sequence
In the encoder–decoder setting, attention allowed the decoder to query the encoder states. But the same mechanism can be applied within a single sequence: self-attention (also called intra-attention) lets every position in a sequence attend to every other position in the same sequence. Q, K, and V are all derived from the same input X.
Self-attention computes a rich, context-aware representation of each token by aggregating information from all other tokens. For the word “bank” in a sentence, self-attention can simultaneously gather evidence from “river” (suggesting a geographic meaning) or “loan” (suggesting a financial meaning), weighting each contextual cue by its relevance to resolving the ambiguity. This context-sensitivity is learned entirely from data — the model develops its own notion of relevance based on the training objective.
Unlike RNN hidden states, which are computed sequentially (position t depends on position t−1), self-attention is computed in parallel across all positions. Every position independently queries all other positions and assembles its context vector. This parallelism is the key reason Transformers train much faster than RNNs on modern hardware, where matrix parallelism is the primary computational resource.
Self-attention has O(n²d) time and space complexity, where n is the sequence length and d is the model dimension. The n×n attention matrix must be computed and stored for each layer and each head. For sequences of length 512, this is manageable. For very long sequences (documents of thousands of tokens, high-resolution images), the quadratic cost becomes a bottleneck. Research into sparse attention, linear attention, and chunked attention (e.g., FlashAttention) addresses this limitation while retaining most of the expressiveness of full attention. The next lesson revisits this quadratic bottleneck when comparing how Transformers and RNNs scale.
Cross-Attention: Attending to Another Sequence
Cross-attention is the generalization that connects two different sequences. Queries come from one sequence (e.g., the decoder), while keys and values come from another (e.g., the encoder). This is exactly the encoder–decoder attention pattern: the decoder queries the encoder representations to extract the information it needs for generating each output token.
Cross-attention generalizes naturally beyond translation. In image captioning, queries from the language decoder attend to visual features extracted by a vision encoder. In speech recognition, queries from the text decoder attend to audio encoder representations. In document question answering, queries from the answer decoder attend to document representations. Wherever two modalities or two sequences must interact, cross-attention provides the mechanism.
The distinction between self-attention and cross-attention is therefore architectural, not mathematical. Both use the same scaled dot-product formula. The difference is only in where Q, K, V come from: same sequence for self-attention, different sequences for cross-attention. This unification is one of the Transformer’s key elegances — a single primitive composes into both intra-sequence and inter-sequence interaction.
A Word on Position
Self-attention is permutation-equivariant: shuffling the input tokens produces the same attention outputs in the same shuffled order. The mechanism has no inherent notion of position — attending to position 1 vs. position 100 looks identical to the attention formula. This is unlike RNNs, where position is implicitly encoded by the order in which hidden states are computed.
To inject positional information, Transformers add positional encodings to the input embeddings before any attention layer. The original Transformer used sinusoidal functions of different frequencies to produce a unique encoding for each position. Modern models often use learned positional embeddings or relative positional encodings that directly bias attention scores based on the distance between positions. We will examine positional encoding in detail in Lesson 9.2 when we study the full Transformer architecture.
- Attention overcomes the fixed-vector bottleneck of encoder–decoder RNNs by giving the decoder direct, weighted access to all encoder hidden states at each decoding step. Weights are computed by a learnable alignment function and normalized with softmax.
- The Query–Key–Value abstraction reformulates attention as differentiable information retrieval: queries ask what to look for, keys describe what each position contains, and values carry the actual information. All three are linear projections of the input.
- Scaled dot-product attention computes all query–key similarities as a single matrix multiply and divides by √dk to prevent gradient saturation. Masking enforces causal constraints or ignores padding.
- Multi-head attention runs h independent attention heads in parallel, each projecting into a lower-dimensional subspace. Heads learn to specialize in different relationship types (syntactic, semantic, positional). Their outputs are concatenated and projected back to dmodel.
- Self-attention applies the same mechanism within a single sequence, letting every position attend to every other position simultaneously and in parallel — unlike RNNs, which process positions sequentially. This parallelism is the key efficiency advantage of Transformers.
- Cross-attention connects two sequences by routing queries from one and keys/values from another. It is the same mathematical operation as self-attention, enabling a unified framework for encoder–decoder interaction, multimodal fusion, and beyond.