ML 101
M08 · L01
Module 8

Sequence Modeling

Text, time series, audio, video — all sequential. Standard networks forget time. Recurrent networks carry a hidden state that threads memory through every step.

01 / 11
ML 101
M08 · L01
Sequential Data

Order Matters

Examples
Text & NLP
Examples
Time Series
Examples
Audio & Video

Shuffling the elements destroys meaning. Models must respect position and capture long-range dependencies across arbitrary distances.

02 / 11
ML 101
M08 · L01
Why Feedforward Fails

Fixed Size, No Memory

  • Requires fixed-size input — sequences vary in length
  • No parameter sharing across time steps — every position has independent weights
  • Sliding window fixes length but hard-codes context — long-range deps invisible
  • Must relearn the same pattern at every position independently
03 / 11
ML 101
M08 · L01
Recurrent Update Rule

Hidden State Memory

At every step t, the same weight matrices combine the new input with the previous hidden state — parameter sharing across time for free.

Simple RNN
h_t = \tanh(W_h h_{t-1} + W_x x_t + b)
04 / 11
ML 101
M08 · L01
Unfolding Through Time

Deep in Time

  • Rewrite recurrence as an acyclic graph: one copy per time step
  • Weights Wh, Wx are tied across all copies — same parameters everywhere
  • Standard backpropagation applies to the unrolled graph
  • Total loss = sum of per-step losses; gradients accumulate from all steps
05 / 11
ML 101
M08 · L01
Backprop Through Time

BPTT & Truncation

The gradient of the total loss w.r.t. Wh sums over all T steps. Full BPTT is expensive for long sequences — truncated BPTT limits propagation to k steps per chunk.

Gradient Accumulation
\frac{\partial L}{\partial W_h} = \sum_{t=1}^{T} \frac{\partial L_t}{\partial W_h}
06 / 11
ML 101
M08 · L01
Vanishing Gradient

Gradient Disappears

  • Gradient through k steps = product of k Jacobians related to Wh
  • Spectral radius < 1 → exponential shrinkage — distant past invisible
  • Spectral radius > 1 → exponential growth — exploding gradients
  • Gradient clipping tames explosions; vanishing requires architecture changes
07 / 11
ML 101
M08 · L01
Practical Remedies

Taming the Gradient

  • Gradient clipping — rescale gradient if norm exceeds threshold θ
  • Orthogonal / identity init of Wh — preserves gradient magnitude early in training
  • Short sequence training — start on short sequences, gradually increase length
  • LSTM / GRU — gating mechanisms that maintain gradient highways (Lesson 8.2)
08 / 11
ML 101
M08 · L01
RNN Topologies

Input–Output Shapes

  • One-to-many — image → caption (image captioning)
  • Many-to-one — sentence → label (sentiment classification)
  • Many-to-many — sequence → sequence (tagging, language modeling)
  • Encoder–Decoder — compress input into context vector; decoder generates output
09 / 11
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 11
ML 101
Summary
Recap

What You Learned

RNNs carry a hidden state updated at every step via ht = tanh(Whht−1 + Wxxt + b). BPTT trains them by unrolling through time. Vanishing gradients limit long-range memory — gradient clipping handles explosions, while LSTM and GRU (next lesson) solve vanishing.

11 / 11