ML 101
M08 · L02
Module 8

LSTM & GRU

Simple RNNs forget the distant past. LSTM and GRU solve vanishing gradients with gates that control what to remember, what to write, and what to output — enabling memory across hundreds of steps.

01 / 11
ML 101
M08 · L02
Cell State

The Memory Highway

The LSTM separates short-term memory (hidden state ht) from long-term memory (cell state ct). The cell state flows with minimal modification — a gradient highway that bypasses vanishing.

Short-term
Hidden ht
Long-term
Cell ct
02 / 11
ML 101
M08 · L02
Three Gates

Control Information Flow

  • Forget gate ft — what fraction of old cell state to erase
  • Input gate it — how much new candidate memory to write
  • Output gate ot — what portion of cell state to expose as ht
  • All gates: sigmoid σ output in (0, 1) — 0 = block, 1 = pass fully
03 / 11
ML 101
M08 · L02
LSTM Update

Additive Cell Update

The cell state update is additive: old memory scaled by forget gate, plus new candidate memory scaled by input gate. Additive updates = gradient highways.

Cell & Hidden Update
\begin{aligned}c_t &= f_t \odot c_{t-1} + i_t \odot g_t\\ h_t &= o_t \odot \tanh(c_t)\end{aligned}
04 / 11
ML 101
M08 · L02
Why It Works

Constant Error Carousel

  • Gradient of loss w.r.t. ct−1 is multiplied only by forget gate ft
  • Network learns ft ≈ 1 for content to preserve — gradient flows unchanged
  • No tanh chain across time steps on the cell highway
  • Enables learning across hundreds of steps — impossible for simple RNN
05 / 11
ML 101
M08 · L02
Gated Recurrent Unit

GRU: Two Gates

GRU merges cell and hidden state into one vector and replaces three LSTM gates with two: reset gate rt (how much past to forget) and update gate zt (how much to carry forward).

GRU Hidden Update
h_t = (1-z_t)\odot h_{t-1}+z_t\odot\tilde{h}_t
06 / 11
ML 101
M08 · L02
LSTM vs GRU

Choosing Between Them

  • GRU has fewer parameters — faster to train, less data required
  • LSTM may edge ahead on very long-range memory tasks
  • Empirically comparable on most NLP / time-series benchmarks
  • Both vastly outperform simple RNN — choice is a hyperparameter
07 / 11
ML 101
M08 · L02
Bidirectional RNNs

Past & Future Context

  • Forward pass (left → right) + backward pass (right → left)
  • Concatenate both hidden states at each position: ht = [&overrightarrow;ht ; &overleftarrow;ht]
  • Every position sees full sequence context — critical for NER, translation
  • Constraint: full sequence must be available — cannot stream in real time
08 / 11
ML 101
M08 · L02
Deep RNNs

Stacking Layers

  • Layer l output ht(l) feeds as input to layer l+1 at same time step
  • 2–4 layers typical — deeper rarely helps for sequences
  • Dropout between layers (not recurrent connections) for regularization
  • Deep BiLSTM was state-of-the-art until Transformers (2017)
09 / 11
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 11
ML 101
Summary
Recap

What You Learned

LSTM solves vanishing gradients with a cell state and three gates. GRU simplifies to two gates with comparable performance. Bidirectional RNNs add future context. Stack 2–4 layers for richer representations. Transformers (Module 9) replace recurrence with parallel self-attention.

11 / 11