Reading
Stories Mode

Sequence Applications

~22 min read Lesson 3 of 3 in Module 8

From Architecture to Application

The previous two lessons built the machinery: recurrent connections that carry memory across time steps, and gating mechanisms that protect long-range dependencies from vanishing gradients. This lesson puts that machinery to work. RNNs, LSTMs, and GRUs are not just theoretical constructs — they powered a decade of breakthroughs across language, time series, speech, and sequential decision-making. Understanding how the architecture maps onto each problem is the key skill for applying sequence models in practice.

We will examine five concrete application domains: language modeling, text classification, sequence-to-sequence translation, time series forecasting, and speech recognition. Each domain illustrates a different way of framing a sequential problem — different input and output structures, different loss functions, different inference procedures. We close by previewing the attention mechanism, the key innovation that eventually led to Transformers and superseded pure recurrence for many tasks.

Language Modeling: Predicting the Next Word

A language model assigns a probability to every sequence of tokens. The standard formulation uses the chain rule of probability: P(w1, w2, …, wT) = ∏t=1T P(wt | w1, …, wt−1). Training reduces to a next-token prediction task: given the sequence so far, predict the next token. An RNN is a natural fit because the hidden state ht−1 compresses all preceding tokens into a fixed-size vector, and a softmax output layer converts that representation into a probability distribution over the vocabulary.

At each time step, the cross-entropy loss measures the discrepancy between the model’s predicted distribution and the true next token (a one-hot vector). The loss is averaged over all positions and all sequences in the mini-batch. During inference, tokens are sampled from the predicted distribution autoregressively: the sampled token becomes the next input, and the process repeats until an end-of-sequence token is generated or the maximum length is reached.

Language Model Loss
\mathcal{L} = -\frac{1}{T}\sum_{t=1}^{T}\log P(w_t \mid w_1,\ldots,w_{t-1})
The language model loss is the average negative log-likelihood of each true token given its prefix. Minimizing this loss trains the network to assign high probability to real continuations. The exponentiated loss is called perplexity — a lower perplexity means the model is less surprised by the data.

Character-level language models tokenize text into individual characters, enabling the model to generate any string and handle out-of-vocabulary words. Word-level models tokenize into words or subwords (e.g., byte-pair encoding), producing more semantically coherent outputs at the cost of a larger vocabulary. Character-level RNNs were among the first compelling demonstrations that neural sequence models could generate coherent prose, code, and structured data.

Perplexity as a Metric

Perplexity (PPL) is the standard metric for language models: PPL = exp(loss). A perplexity of k means the model is, on average, as confused as if it had to choose uniformly among k equally likely options at each step. A random model over a 50,000-word vocabulary has perplexity 50,000. A well-trained word-level RNN typically achieves perplexity in the range of 60–100 on the Penn Treebank benchmark. Modern Transformer language models push below 20 on the same benchmark.

Text Classification: Encoding a Sequence into a Label

Text classification maps a variable-length sequence to a fixed categorical label: positive or negative sentiment, spam or not spam, the topic of an article. The natural RNN formulation is to run the sequence through the recurrent network and use the final hidden state hT as a summary representation of the entire sequence. A fully connected classification head on top of hT maps this fixed-size vector to class probabilities via softmax.

The key insight is that hT is a learned summary of the full input sequence, compressed through the recurrent transitions and gating mechanisms. For an LSTM, the cell state cT often provides a complementary signal and is concatenated with hT before the classification head. For bidirectional models, both final hidden states [&overrightarrow;hT ; &overleftarrow;h1] are concatenated, capturing both forward and backward context.

An important practical issue is handling variable-length sequences in a mini-batch. Sequences are padded to the same length with a special PAD token, and a masking mechanism ensures that the loss is computed only on real tokens, not padding. The final hidden state is taken from the last real (non-padded) position for each sequence, not from the padded end. Modern frameworks provide efficient packed-sequence representations that avoid computing over padding entirely.

Sequence-to-Sequence: Translation and Beyond

Many tasks require generating a variable-length output sequence from a variable-length input sequence: machine translation (English → French), summarization (long document → short summary), dialogue (user utterance → system response). These tasks cannot be solved by a single RNN because the output length is not determined by the input length.

The encoder–decoder architecture (Sutskever et al., 2014) addresses this by splitting the model into two RNNs. The encoder reads the full input sequence and compresses it into a fixed-size context vector c, typically the final hidden state. The decoder is initialized with c and autoregressively generates the output sequence one token at a time, feeding each generated token as input to the next step. The decoder continues until it generates a special end-of-sequence token.

Encoder–Decoder
\begin{aligned}c &= h_T^{\text{enc}}\\ s_t &= f_{\text{dec}}(s_{t-1},y_{t-1},c)\\ P(y_t) &= \text{softmax}(W_o s_t)\end{aligned}
The encoder compresses the input into context c = hTenc. The decoder generates output tokens yt conditioned on the context c, its previous hidden state st−1, and the previous output token yt−1. All parameters are trained jointly by maximizing the log-probability of the correct output sequence.

The bottleneck of this architecture is the single context vector: the entire input sequence — a sentence of 20 or 50 words — must be compressed into a single fixed-size vector. For long sequences, the encoder’s last hidden state simply cannot retain all relevant information. This bottleneck motivated the attention mechanism, which allows the decoder to look back at all encoder hidden states rather than relying on a single summary.

Beam Search at Inference Time

Greedy decoding at each step — always picking the most probable next token — often produces suboptimal output sequences, because an early greedy choice can preclude better completions. Beam search maintains the top-k candidate sequences (the beam) at each step, expanding all candidates and keeping only the k most probable partial sequences. A beam width of 4–10 substantially improves translation quality over greedy decoding, at the cost of k-fold inference compute. Length normalization (dividing log-probability by sequence length) prevents beam search from favoring shorter outputs.

Time Series Forecasting

Beyond language, RNNs were applied to numerical time series: predicting future values of a signal given its history. A stock price one day ahead, the electricity demand in the next hour, the remaining useful life of an industrial machine — all are forecasting tasks that map a sequence of past observations to one or more future values.

For one-step-ahead forecasting, the model takes x1, …, xT as input and predicts xT+1. The output is a continuous value, so the loss is mean squared error (MSE) or mean absolute error (MAE) rather than cross-entropy. The architecture is essentially the same as an RNN classifier, but the output layer is a linear regression instead of a softmax. For multi-step forecasting, the decoder can generate future values autoregressively, or a single output layer can simultaneously predict all H future steps (the direct strategy).

A key challenge in time series is non-stationarity: the statistical properties of the series (mean, variance, seasonality) change over time. Standard preprocessing includes differencing (subtracting the previous value to remove trends), normalization (zero mean, unit variance), and seasonal decomposition. LSTMs have been especially effective for multivariate time series, where correlations among multiple signals (e.g., temperature, humidity, and wind speed predicting energy demand) must be captured jointly.

Speech Recognition

Automatic speech recognition (ASR) maps a sequence of acoustic feature frames (typically 25ms log-mel filterbank vectors, extracted every 10ms) to a sequence of characters or words. The input and output sequences have different lengths — a 3-second utterance might produce 300 acoustic frames and 40 characters — and the alignment between frames and output tokens is unknown at training time.

The Connectionist Temporal Classification (CTC) loss, introduced by Graves et al. in 2006, solved the alignment problem without requiring hand-labeled frame-level transcripts. CTC introduces a special blank token and marginalizes over all valid alignments between the output sequence and the input frames, computing the total probability of the correct output sequence over all possible frame-to-token assignments. This allows training from text transcripts alone, with the model learning the alignment implicitly.

CTC Loss
\mathcal{L}_{\text{CTC}} = -\log\sum_{\pi\in\mathcal{B}^{-1}(y)} P(\pi \mid x)
The CTC loss maximizes the total probability of all valid alignments π that collapse (by merging repeated tokens and removing blanks) to the correct transcription y. Β−1(y) denotes the set of all valid CTC paths for output y. Efficient computation uses a forward-backward dynamic programming algorithm.

Deep bidirectional LSTMs trained with CTC achieved near-human word error rates on clean-speech benchmarks by 2015. The standard architecture stacked 5–7 bidirectional LSTM layers on top of acoustic features, with batch normalization and dropout for regularization. ASR was one of the first real-world applications to demonstrate that deep recurrent networks, when trained on thousands of hours of audio, could match and exceed human-level accuracy on narrow tasks.

Attention: A Preview

The encoder–decoder bottleneck and the inherent sequentiality of RNNs both pointed to the same limitation: important information might be too far away in the hidden-state chain to be retrieved reliably. The attention mechanism, introduced by Bahdanau et al. in 2015 for neural machine translation, addressed this directly.

Instead of compressing the input into a single context vector, the attention mechanism allows the decoder to compute a weighted sum over all encoder hidden states at each decoding step. The weights (attention scores) are computed by a small learned alignment network that compares the current decoder state st with each encoder state hj. The resulting context vector ct is a soft, differentiable lookup into the encoder’s memory, focused on the parts of the input most relevant to producing the current output token.

Bahdanau Attention
\begin{aligned}e_{tj} &= a(s_{t-1},h_j)\\ \alpha_{tj} &= \frac{\exp(e_{tj})}{\sum_k \exp(e_{tk})}\\ c_t &= \sum_j \alpha_{tj}h_j\end{aligned}
The alignment score etj measures how well the decoder state st matches encoder state hj. Softmax normalizes scores into attention weights αtj that sum to 1. The context vector ct is the weighted sum of encoder states. The decoder then generates yt from (st, ct).

Attention dramatically improved translation quality on long sentences, where the fixed context vector was insufficient. Visualizing the attention weight matrix produces interpretable alignment patterns — for English-to-French translation, the model often attends to the correct source word when generating each target word, even when the word order differs. Attention also enabled the model to handle sentences of arbitrary length without degradation, because the decoder always has direct access to all encoder hidden states regardless of sequence length.

This “soft lookup” idea was then generalized: what if attention was not just an add-on to RNNs, but the primary mechanism? The Transformer architecture (Module 9) removes recurrence entirely, replacing it with self-attention — each position attending to all other positions in the same sequence simultaneously. The result is a model that can be parallelized during training (since there is no sequential dependency across positions), scales to much longer sequences, and has become the foundation of virtually every modern large language model.

When to Use RNNs Today

Transformers have largely superseded RNNs for tasks where the full sequence is available (translation, NLP, time series analysis on fixed windows). But RNNs retain practical advantages for streaming inference: processing one step at a time with O(1) compute per step and bounded memory. An IoT sensor that needs to classify events in real time, a speech synthesis system that must generate audio samples continuously, or a robot controller processing sensor streams — all benefit from the RNN’s sequential efficiency. When data is limited or compute is constrained, the smaller parameter count of GRUs and LSTMs is also an advantage over large Transformer models.

Key Takeaways
  • Language models are trained with next-token prediction (cross-entropy loss) and evaluated with perplexity. The RNN hidden state compresses the prefix into a fixed-size vector, from which a softmax output produces token probabilities. Autoregressive sampling generates new sequences.
  • Text classification uses the final hidden state (or both final states in bidirectional RNNs) as a fixed-size sequence representation, fed into a classification head. Variable-length sequences are handled by padding and masking. The loss is cross-entropy over class labels.
  • Sequence-to-sequence models use an encoder RNN to compress the input into a context vector and a decoder RNN to generate the output autoregressively. The single-vector bottleneck limits performance on long sequences. Beam search improves inference quality over greedy decoding.
  • Time series forecasting treats numerical signals as sequences, using MSE/MAE loss and linear output layers instead of softmax. LSTMs capture multi-step temporal dependencies and cross-variable correlations effectively, with preprocessing (differencing, normalization) handling non-stationarity.
  • Speech recognition maps acoustic frame sequences to characters using CTC loss, which marginalizes over all valid frame-to-token alignments. Deep bidirectional LSTMs trained with CTC achieved near-human word error rates by 2015.
  • The attention mechanism allows the decoder to compute a weighted sum over all encoder states at each step, overcoming the context vector bottleneck. Generalizing attention to self-attention — where every position attends to every other position — led directly to the Transformer architecture covered in Module 9.
Previous LSTM and GRU Overview Next The Attention Mechanism