LSTM & GRU
Simple RNNs forget the distant past. LSTM and GRU solve vanishing gradients with gates that control what to remember, what to write, and what to output — enabling memory across hundreds of steps.
The Memory Highway
The LSTM separates short-term memory (hidden state ht) from long-term memory (cell state ct). The cell state flows with minimal modification — a gradient highway that bypasses vanishing.
Control Information Flow
- Forget gate ft — what fraction of old cell state to erase
- Input gate it — how much new candidate memory to write
- Output gate ot — what portion of cell state to expose as ht
- All gates: sigmoid σ output in (0, 1) — 0 = block, 1 = pass fully
Additive Cell Update
The cell state update is additive: old memory scaled by forget gate, plus new candidate memory scaled by input gate. Additive updates = gradient highways.
Constant Error Carousel
- Gradient of loss w.r.t. ct−1 is multiplied only by forget gate ft
- Network learns ft ≈ 1 for content to preserve — gradient flows unchanged
- No tanh chain across time steps on the cell highway
- Enables learning across hundreds of steps — impossible for simple RNN
GRU: Two Gates
GRU merges cell and hidden state into one vector and replaces three LSTM gates with two: reset gate rt (how much past to forget) and update gate zt (how much to carry forward).
Choosing Between Them
- GRU has fewer parameters — faster to train, less data required
- LSTM may edge ahead on very long-range memory tasks
- Empirically comparable on most NLP / time-series benchmarks
- Both vastly outperform simple RNN — choice is a hyperparameter
Past & Future Context
- Forward pass (left → right) + backward pass (right → left)
- Concatenate both hidden states at each position: ht = [&overrightarrow;ht ; &overleftarrow;ht]
- Every position sees full sequence context — critical for NER, translation
- Constraint: full sequence must be available — cannot stream in real time
Stacking Layers
- Layer l output ht(l) feeds as input to layer l+1 at same time step
- 2–4 layers typical — deeper rarely helps for sequences
- Dropout between layers (not recurrent connections) for regularization
- Deep BiLSTM was state-of-the-art until Transformers (2017)
What You Learned
LSTM solves vanishing gradients with a cell state and three gates. GRU simplifies to two gates with comparable performance. Bidirectional RNNs add future context. Stack 2–4 layers for richer representations. Transformers (Module 9) replace recurrence with parallel self-attention.