Sequence Applications
Language models, text classifiers, seq2seq translators, time series forecasters, and speech recognizers — every major RNN application mapped to a concrete architecture, loss function, and inference strategy.
Predict the Next Word
Train by next-token prediction: given prefix w1…wt−1, predict wt. Loss = cross-entropy. Evaluation metric = perplexity = exp(loss). Lower perplexity means the model is less surprised by real text.
Language Model Loss
Minimize average negative log-likelihood over all positions. Autoregressive sampling at inference: generate token → feed back as next input → repeat.
Sequence to Label
- Run sequence through RNN → use final hidden state hT as summary
- Add linear classification head + softmax for class probabilities
- Bidirectional: concatenate [&overrightarrow;hT ; &overleftarrow;h1] for richer representation
- Padding + masking handles variable-length sequences in mini-batches
Encoder — Decoder
- Encoder RNN compresses input into context vector c = hTenc
- Decoder RNN initialized with c, generates output autoregressively
- Applications: translation, summarization, dialogue
- Bottleneck: whole input compressed into one fixed-size vector
Beam Search
Greedy decoding picks the most probable token at each step but can miss better overall sequences. Beam search maintains the top-k candidates at each step — beam width 4–10 substantially improves quality.
Predict Future Values
- Output layer: linear regression (MSE/MAE loss), not softmax
- One-step-ahead or multi-step direct/autoregressive strategies
- Preprocessing: differencing, normalization for non-stationarity
- LSTM captures cross-variable correlations in multivariate series
Acoustic Frames to Text
25ms log-mel filterbank frames → deep BiLSTM → characters. Input and output have different lengths. CTC loss marginalizes over all valid frame-to-token alignments — no hand-labeled frame alignment needed.
Soft Lookup
- Decoder computes weighted sum over all encoder states at each step
- Alignment scores: a(st−1, hj) → softmax → attention weights αtj
- Context ct = ∑jαtjhj focuses on relevant input positions
- Overcomes fixed-vector bottleneck — led directly to Transformers
What You Learned
RNNs power language models (next-token prediction + perplexity), text classifiers (final hidden state + softmax), seq2seq translation (encoder–decoder + beam search), time series forecasting (MSE loss + linear output), and speech recognition (CTC loss). Attention overcomes the context-vector bottleneck, paving the way for Transformers in Module 9.