DSP 101
M9 · L3
Module 9 — Spectral Analysis
Spectrogram — Time-Frequency Analysis
Beyond the PSD — how to reveal how a signal's frequency content changes over time.
1 / 9
DSP 101
M9 · L3
The Problem with PSD
When Signals Change
The PSD assumes stationarity — that spectral content is constant in time. Most real signals break this assumption.
- Speech: silence → vowels → consonants — all in milliseconds
- Radar: bat chirps sweep 80 kHz → 20 kHz in <1 ms
- Seismic: quiet until the P-wave arrives
- Music: notes start, sustain, and decay in time
- We need: frequency content as a function of time
2 / 9
DSP 101
M9 · L3
The Core Tool
Short-Time Fourier Transform
Slide a short window along the signal and compute a DFT at each position. Each windowed DFT gives local frequency content.
STFT
X(m, \omega) = \sum_{n} x[n]\, w[n - mR]\, e^{-j\omega n}
Spectrogram
S(m,ω) = |X(m,ω)|² — squared magnitude of STFT. A 2-D time-frequency power map.
3 / 9
DSP 101
M9 · L3
The Fundamental Limit
Time vs. Frequency Resolution
Window length M sets both resolutions — you cannot improve one without degrading the other.
↓M
Fine time, coarse freq
↑M
Fine freq, coarse time
Uncertainty principle
Δt · Δf ≥ 1/(4π) — no window escapes this bound. This is the signal-processing version of Heisenberg's principle.
4 / 9
DSP 101
M9 · L3
Design Choices
STFT Parameters
- Window length M: controls the core resolution trade-off
- Hop size R: M/4 or M/2 — smaller R → more frames (denser time grid, not higher resolution)
- Window type: Hann is standard — reduces spectral leakage between frames
- FFT size N ≥ M: zero-padding gives a finer frequency grid but no new information
- Rule of thumb for speech/audio: M ≈ 20–25 ms at the sample rate
5 / 9
DSP 101
M9 · L3
Visual Interpretation
Reading Spectrograms
Time → right, frequency → up, color → power (dB). Common features:
- Horizontal line: constant-frequency tone or sinusoid
- Diagonal line: frequency chirp (FM sweep, bat echolocation)
- Harmonic stack: f₀, 2f₀, 3f₀ — voiced speech, musical instruments
- Vertical stripe: broadband burst (click, drum hit)
- Diffuse cloud: noise or sibilance (/s/, /sh/)
- Formant tracks: slowly varying bands — vowel resonances
6 / 9
DSP 101
M9 · L3
The Mel Spectrogram & Beyond
Where Spectrograms Matter
- Mel spectrogram: STFT → mel filterbank → log = perceptual features for ML (Whisper, music AI)
- Radar / Sonar: micro-Doppler, propeller harmonics, target classification
- Vibration: bearing fault detection at characteristic fault frequencies
- EEG: seizure detection — sudden broadband power increase
- SIGINT: identify signal type, modulation, bandwidth from spectrogram
- Seismology: P-wave / S-wave arrival times, rupture frequency content
7 / 9
DSP 101
M9 · L3
Key Takeaways
What You Learned
- Spectrogram = |STFT|² — a 2-D time-frequency power map for non-stationary signals
- STFT: slide a windowed DFT over time with hop size R
- Short window → fine time, coarse frequency; long window → reverse
- Δt · Δf ≥ 1/(4π) — the uncertainty principle is inescapable
- Hop size and zero-padding affect grid density, not true resolution
- Mel spectrogram = STFT → mel filterbank → log — the standard audio ML feature
- Phase is discarded — reconstruction needs iterative algorithms (Griffin-Lim)
9 / 9
1 / 8
muted