The Attention Mechanism
From fixed-vector bottleneck to dynamic lookup — how attention gives every decoder step direct access to any encoder position, and why this single idea reshaped all of deep learning.
The Bottleneck
Encoder–decoder RNNs compress the entire input into one fixed-size vector. For long sequences, critical information is lost. Translation quality degrades with sentence length. Attention discards this bottleneck entirely.
Weighted Retrieval
Alignment scores etj measure relevance. Softmax gives weights αtj summing to 1. Context ct = weighted sum of all encoder states. Fully differentiable — no explicit alignment supervision needed.
Query · Key · Value
- Query (Q) — what am I looking for?
- Key (K) — what does each position contain?
- Value (V) — what information does each position return?
- All three are linear projections of the input: Q=XWQ, K=XWK, V=XWV
One Matrix Multiply
Divide by √dk to prevent gradient saturation. Optional mask sets future positions to −∞ for causal attention.
Variance Control
Random dot products between unit-normal Q and K vectors have variance dk. Without scaling, large dk makes softmax nearly one-hot — near-zero gradients, slow training.
h Heads in Parallel
- h independent attention heads, each in dk = dmodel/h dimensions
- Each head learns a different type of relationship (syntactic, semantic, positional)
- Outputs concatenated and projected back to dmodel by WO
- Same total parameters as a single full-dimension head
Every Position Attends to Every Other
- Q, K, V all derived from the same input sequence X
- Each token gathers context from all other tokens simultaneously
- Computed in parallel — no sequential dependency like RNNs
- Cost: O(n²d) — quadratic in sequence length
Two Sequences Interact
- Queries from one sequence (decoder), Keys & Values from another (encoder)
- Same math as self-attention — only source of Q, K, V differs
- Enables: translation, image captioning, speech recognition
- A single primitive for all cross-modal and cross-sequence interaction
What You Learned
Attention replaces the fixed-vector bottleneck with dynamic weighted retrieval. The Q–K–V abstraction unifies Bahdanau attention, scaled dot-product attention, multi-head attention, self-attention, and cross-attention into one elegant primitive. Module 9.2 assembles these pieces into the full Transformer architecture.