ML 101
M08 · L03
Module 8

Sequence Applications

Language models, text classifiers, seq2seq translators, time series forecasters, and speech recognizers — every major RNN application mapped to a concrete architecture, loss function, and inference strategy.

01 / 11
ML 101
M08 · L03
Language Modeling

Predict the Next Word

Train by next-token prediction: given prefix w1…wt−1, predict wt. Loss = cross-entropy. Evaluation metric = perplexity = exp(loss). Lower perplexity means the model is less surprised by real text.

RNN LM PPL
60–100
Transformer PPL
< 20
02 / 11
ML 101
M08 · L03
LM Training

Language Model Loss

Cross-Entropy Loss
-\frac{1}{T}\sum_{t=1}^{T}\log P(w_t\mid w_1,\ldots,w_{t-1})

Minimize average negative log-likelihood over all positions. Autoregressive sampling at inference: generate token → feed back as next input → repeat.

03 / 11
ML 101
M08 · L03
Text Classification

Sequence to Label

  • Run sequence through RNN → use final hidden state hT as summary
  • Add linear classification head + softmax for class probabilities
  • Bidirectional: concatenate [&overrightarrow;hT ; &overleftarrow;h1] for richer representation
  • Padding + masking handles variable-length sequences in mini-batches
04 / 11
ML 101
M08 · L03
Sequence-to-Sequence

Encoder — Decoder

  • Encoder RNN compresses input into context vector c = hTenc
  • Decoder RNN initialized with c, generates output autoregressively
  • Applications: translation, summarization, dialogue
  • Bottleneck: whole input compressed into one fixed-size vector
05 / 11
ML 101
M08 · L03
Inference

Beam Search

Greedy decoding picks the most probable token at each step but can miss better overall sequences. Beam search maintains the top-k candidates at each step — beam width 4–10 substantially improves quality.

Greedy
k = 1
Beam
k = 4–10
06 / 11
ML 101
M08 · L03
Time Series Forecasting

Predict Future Values

  • Output layer: linear regression (MSE/MAE loss), not softmax
  • One-step-ahead or multi-step direct/autoregressive strategies
  • Preprocessing: differencing, normalization for non-stationarity
  • LSTM captures cross-variable correlations in multivariate series
07 / 11
ML 101
M08 · L03
Speech Recognition

Acoustic Frames to Text

25ms log-mel filterbank frames → deep BiLSTM → characters. Input and output have different lengths. CTC loss marginalizes over all valid frame-to-token alignments — no hand-labeled frame alignment needed.

CTC Loss
-\log\sum_{\pi\in\mathcal{B}^{-1}(y)}P(\pi\mid x)
08 / 11
ML 101
M08 · L03
Attention Mechanism

Soft Lookup

  • Decoder computes weighted sum over all encoder states at each step
  • Alignment scores: a(st−1, hj) → softmax → attention weights αtj
  • Context ct = ∑jαtjhj focuses on relevant input positions
  • Overcomes fixed-vector bottleneck — led directly to Transformers
09 / 11
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 11
ML 101
Summary
Recap

What You Learned

RNNs power language models (next-token prediction + perplexity), text classifiers (final hidden state + softmax), seq2seq translation (encoder–decoder + beam search), time series forecasting (MSE loss + linear output), and speech recognition (CTC loss). Attention overcomes the context-vector bottleneck, paving the way for Transformers in Module 9.

11 / 11