The Limitation of the PSD
Both the periodogram and Welch's method produce a single PSD estimate — a function of frequency alone. They assume the signal is stationary: that its statistical properties do not change over time. For many real signals, however, this assumption breaks down completely. Speech transitions between silence, vowels, and consonants in milliseconds. A bat's echolocation call sweeps from 80 kHz to 20 kHz in a fraction of a second. A seismic signal is quiet until the P-wave arrives, then suddenly filled with energy.
For these non-stationary signals, we need a representation that captures how the frequency content changes over time — a joint time-frequency distribution. The spectrogram is the most widely used such representation.
The Short-Time Fourier Transform (STFT)
The spectrogram is computed from the Short-Time Fourier Transform (STFT). The idea is simple: instead of computing one DFT over the entire signal, slide a short analysis window along the signal and compute the DFT at each position. Each windowed DFT gives the local frequency content near that time position.
The spectrogram is simply the squared magnitude of the STFT — the time-frequency power distribution:
The Time-Frequency Resolution Trade-Off
The most fundamental property of the spectrogram is its time-frequency resolution trade-off. This is an unavoidable consequence of how windowing works in the Fourier transform, and it is the discrete-signal analog of the Heisenberg uncertainty principle in quantum mechanics.
The duration of the analysis window M determines both resolutions simultaneously:
- Fine time resolution — locates events precisely in time
- Poor frequency resolution — Δf = f_s/M is large
- Blurry, broad peaks in frequency
- Good for transients: clicks, onsets, impacts
- Typical: M = 128–512 samples
- Poor time resolution — smears events in time
- Fine frequency resolution — Δf = f_s/M is small
- Sharp, narrow peaks in frequency
- Good for harmonics: pitched sounds, tones
- Typical: M = 2048–8192 samples
The product of time resolution (Δt) and frequency resolution (Δf) is bounded below by a constant that depends on the window shape. For a Gaussian window it achieves the theoretical minimum (the Gabor limit). No window can give arbitrarily fine resolution in both dimensions simultaneously — you can shift the trade-off with window choice, but you cannot escape it.
Δt · Δf ≥ 1/(4π) — the Gabor limit for any linear time-frequency analysisSTFT Parameters in Practice
The spectrogram is controlled by four parameters. Each affects the time-frequency resolution, computational cost, and visual appearance of the result:
| Parameter | Effect on Time Resolution | Effect on Frequency Resolution | Typical Value |
|---|---|---|---|
| Window length M | ↑M → coarser (longer frames) | ↑M → finer (Δf = f_s/M) | 512–4096 samples |
| Hop size R | ↓R → finer (more overlap) | No effect | M/4 or M/2 |
| Window type | Affects sidelobe leakage | Wider main lobe = coarser Δf | Hann (most common) |
| FFT size N | No effect | ↑N → finer frequency grid (zero-padding) | N ≥ M (next power of 2) |
Notice that the hop size R and the FFT size N do not change the fundamental resolution — they only change the sampling density of the time-frequency plane. A smaller hop gives more time frames (interpolated, not truly higher resolution). Zero-padding by choosing N > M gives a finer frequency grid but does not reveal any frequency detail not already present in the M-sample window.
Reading a Spectrogram
A spectrogram is displayed as a 2-D color image. Convention: time increases left to right, frequency increases bottom to top, and color encodes power — typically using a perceptually uniform colormap (viridis, inferno) or a simple grayscale. Nearly always plotted in dB: 10 log₁₀(|X(m,ω)|²).
Key features to look for in a spectrogram:
| Feature | What It Looks Like | Example Signal |
|---|---|---|
| Horizontal line | Constant-frequency sinusoid | Pure tone, tuning fork |
| Diagonal line (chirp) | Frequency sweeping linearly in time | FM sweep, bat echolocation |
| Harmonic stack | Multiple parallel horizontal lines at f₀, 2f₀, 3f₀, … | Voiced speech, musical instruments |
| Vertical stripe | Broadband energy concentrated in time | Click, drum hit, impulse |
| Diffuse cloud | Noise-like energy spread in time and frequency | White noise, frication in speech (/s/, /sh/) |
| Formant tracks | Slowly varying bright bands | Vowel resonances in speech |
The Spectrogram in Speech and Audio
The spectrogram was historically developed for speech analysis and is still the dominant tool in that field. A standard narrowband spectrogram (long window, ~20–30 ms) resolves individual harmonics of the fundamental frequency and tracks formant resonances — the broad spectral peaks that determine vowel identity. A broadband spectrogram (short window, ~3–5 ms) resolves time details like pitch periods and individual glottal pulses but smears harmonics into broad bands.
For audio machine learning (speech recognition, music genre classification, audio tagging), raw spectrograms are rarely used directly. Instead, the frequency axis is mapped to the mel scale — a perceptual scale that mimics how the human auditory system perceives pitch (roughly logarithmic at high frequencies). The mel spectrogram is then compressed using log magnitude, producing a compact, perceptually meaningful feature. This is the input to most modern audio neural networks (CNNs, Transformers, Whisper, etc.).
Mel spectrogram = STFT magnitude → mel filterbank → log compressionSpectrogram Applications Beyond Speech
The spectrogram's ability to reveal how spectral content evolves over time makes it useful across engineering domains:
| Domain | Application | What to Look For |
|---|---|---|
| Radar | Micro-Doppler signature analysis | Frequency modulations from rotating blades, walking humans |
| Sonar | Underwater target classification | Propeller harmonics, cavitation patterns |
| Vibration | Bearing fault detection | Impulsive features at fault frequencies, sidebands |
| Biomedical | EEG seizure detection | Sudden broadband power increase, rhythmic oscillations |
| Communications | Signal identification / SIGINT | Modulation type, bandwidth, burst timing |
| Seismology | Earthquake analysis | P-wave and S-wave arrivals, frequency content of rupture |
Limitations of the Spectrogram
The spectrogram has two fundamental limitations beyond the resolution trade-off. First, it discards phase information — |X(m,ω)|² contains no phase, so you cannot reconstruct the original signal from the spectrogram alone (though approximate reconstruction is possible via iterative algorithms like Griffin-Lim). Second, the spectrogram can exhibit cross-terms for multi-component signals: interference patterns between signal components that appear as spurious features in the time-frequency plane. The Wigner-Ville distribution provides optimal time-frequency concentration but suffers more severely from cross-terms.
For audio at 44.1 kHz, a 20 ms window (882 samples, typically rounded to 1024) gives good harmonic resolution for pitched sounds. For speech at 16 kHz, a 25 ms window (400 samples, rounded to 512) is the standard in speech recognition. For vibration monitoring at 10 kHz, choose a window that gives Δf fine enough to separate the fault frequency from adjacent harmonics. Always verify by visual inspection of the spectrogram.
Rule of thumb: window length ≈ 20–40 ms for speech and audio signals- The spectrogram is the squared magnitude of the STFT — a 2-D time-frequency power map suited for non-stationary signals.
- The STFT slides an analysis window along the signal and computes the DFT at each position with hop size R.
- The fundamental trade-off: short window → fine time resolution, coarse frequency resolution; long window → the reverse. Both cannot be improved simultaneously (uncertainty principle).
- Hop size and zero-padding do not change fundamental resolution — they only affect the density of the time-frequency grid.
- Spectrograms are typically displayed in dB on a color image: time (horizontal), frequency (vertical), power (color).
- The mel spectrogram maps the frequency axis to a perceptual scale and is the standard input feature for audio machine learning systems.
- Phase is discarded in the spectrogram — reconstruction requires iterative phase estimation algorithms.