Sequence Modeling
Text, time series, audio, video — all sequential. Standard networks forget time. Recurrent networks carry a hidden state that threads memory through every step.
Order Matters
Shuffling the elements destroys meaning. Models must respect position and capture long-range dependencies across arbitrary distances.
Fixed Size, No Memory
- Requires fixed-size input — sequences vary in length
- No parameter sharing across time steps — every position has independent weights
- Sliding window fixes length but hard-codes context — long-range deps invisible
- Must relearn the same pattern at every position independently
Hidden State Memory
At every step t, the same weight matrices combine the new input with the previous hidden state — parameter sharing across time for free.
Deep in Time
- Rewrite recurrence as an acyclic graph: one copy per time step
- Weights Wh, Wx are tied across all copies — same parameters everywhere
- Standard backpropagation applies to the unrolled graph
- Total loss = sum of per-step losses; gradients accumulate from all steps
BPTT & Truncation
The gradient of the total loss w.r.t. Wh sums over all T steps. Full BPTT is expensive for long sequences — truncated BPTT limits propagation to k steps per chunk.
Gradient Disappears
- Gradient through k steps = product of k Jacobians related to Wh
- Spectral radius < 1 → exponential shrinkage — distant past invisible
- Spectral radius > 1 → exponential growth — exploding gradients
- Gradient clipping tames explosions; vanishing requires architecture changes
Taming the Gradient
- Gradient clipping — rescale gradient if norm exceeds threshold θ
- Orthogonal / identity init of Wh — preserves gradient magnitude early in training
- Short sequence training — start on short sequences, gradually increase length
- LSTM / GRU — gating mechanisms that maintain gradient highways (Lesson 8.2)
Input–Output Shapes
- One-to-many — image → caption (image captioning)
- Many-to-one — sentence → label (sentiment classification)
- Many-to-many — sequence → sequence (tagging, language modeling)
- Encoder–Decoder — compress input into context vector; decoder generates output
What You Learned
RNNs carry a hidden state updated at every step via ht = tanh(Whht−1 + Wxxt + b). BPTT trains them by unrolling through time. Vanishing gradients limit long-range memory — gradient clipping handles explosions, while LSTM and GRU (next lesson) solve vanishing.