ML 101
M09 · L01
Module 9

The Attention Mechanism

From fixed-vector bottleneck to dynamic lookup — how attention gives every decoder step direct access to any encoder position, and why this single idea reshaped all of deep learning.

01 / 11
ML 101
M09 · L01
The Problem

The Bottleneck

Encoder–decoder RNNs compress the entire input into one fixed-size vector. For long sequences, critical information is lost. Translation quality degrades with sentence length. Attention discards this bottleneck entirely.

Encoder output
N states
RNN bottleneck
1 vector
02 / 11
ML 101
M09 · L01
Bahdanau Attention (2015)

Weighted Retrieval

Context Vector
c_t=\sum_{j=1}^{N}\alpha_{tj}\,h_j,\quad \alpha_{tj}=\frac{\exp(e_{tj})}{\sum_k\exp(e_{tk})}

Alignment scores etj measure relevance. Softmax gives weights αtj summing to 1. Context ct = weighted sum of all encoder states. Fully differentiable — no explicit alignment supervision needed.

03 / 11
ML 101
M09 · L01
The Abstraction

Query · Key · Value

  • Query (Q) — what am I looking for?
  • Key (K) — what does each position contain?
  • Value (V) — what information does each position return?
  • All three are linear projections of the input: Q=XWQ, K=XWK, V=XWV
04 / 11
ML 101
M09 · L01
Scaled Dot-Product Attention

One Matrix Multiply

Attention Formula
\text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V

Divide by √dk to prevent gradient saturation. Optional mask sets future positions to −∞ for causal attention.

05 / 11
ML 101
M09 · L01
Why Scale?

Variance Control

Random dot products between unit-normal Q and K vectors have variance dk. Without scaling, large dk makes softmax nearly one-hot — near-zero gradients, slow training.

Unscaled var
dk
Scaled var
1
06 / 11
ML 101
M09 · L01
Multi-Head Attention

h Heads in Parallel

  • h independent attention heads, each in dk = dmodel/h dimensions
  • Each head learns a different type of relationship (syntactic, semantic, positional)
  • Outputs concatenated and projected back to dmodel by WO
  • Same total parameters as a single full-dimension head
07 / 11
ML 101
M09 · L01
Self-Attention

Every Position Attends to Every Other

  • Q, K, V all derived from the same input sequence X
  • Each token gathers context from all other tokens simultaneously
  • Computed in parallel — no sequential dependency like RNNs
  • Cost: O(n²d) — quadratic in sequence length
08 / 11
ML 101
M09 · L01
Cross-Attention

Two Sequences Interact

  • Queries from one sequence (decoder), Keys & Values from another (encoder)
  • Same math as self-attention — only source of Q, K, V differs
  • Enables: translation, image captioning, speech recognition
  • A single primitive for all cross-modal and cross-sequence interaction
09 / 11
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 11
ML 101
Summary
Recap

What You Learned

Attention replaces the fixed-vector bottleneck with dynamic weighted retrieval. The Q–K–V abstraction unifies Bahdanau attention, scaled dot-product attention, multi-head attention, self-attention, and cross-attention into one elegant primitive. Module 9.2 assembles these pieces into the full Transformer architecture.

11 / 11