The Periodogram — Revisited
In the previous lesson we introduced the periodogram as the simplest estimator of the Power Spectral Density (PSD). Given N samples x[0], x[1], …, x[N−1], the periodogram computes the squared magnitude of the DFT, normalized by N:
The periodogram is appealingly simple, but it suffers from a fundamental statistical problem: its variance does not decrease as the data length N increases. In fact, the variance of each spectral estimate is approximately equal to the square of the true PSD value — roughly 100% relative error at every frequency. No matter how much data you collect, the periodogram remains a noisy estimate.
To understand why, consider that collecting more data with a single DFT gives you finer frequency resolution (smaller Δf = f_s/N), but each frequency bin still corresponds to only one complex-valued DFT sample. The "sample size" per bin never increases — you're always computing one noisy measurement per frequency.
Bias and Variance of the Periodogram
Statistical estimators are characterized by two properties: bias (systematic error) and variance (random fluctuation). For the periodogram:
| Property | Behavior | Effect |
|---|---|---|
| Bias | Decreases as N → ∞ | Asymptotically unbiased — approaches true PSD |
| Variance | ≈ [S(ω)]² — constant! | Does NOT decrease with more data |
| Consistency | Not consistent | Never converges to the true PSD |
The periodogram is therefore an inconsistent estimator — in statistics, an estimator that does not converge to the true value as sample size grows. This is a critical limitation. For reliable PSD estimates, we need methods that reduce variance by averaging multiple independent spectral estimates.
Each DFT bin X[k] involves contributions from all N samples — so it's not "one sample" in an intuitive sense. But mathematically, each |X[k]|² follows a scaled chi-squared distribution with 2 degrees of freedom, regardless of N. Averaging chi-squared variables reduces variance — and that's exactly what Bartlett's and Welch's methods do.
Var[periodogram] ≈ S²(ω) — independent of NBartlett's Method — Averaging Non-Overlapping Segments
The key insight to reducing periodogram variance is averaging. If you could compute the periodogram independently on K different data records and average the results, the variance would decrease by a factor of K. Bartlett's method (1948) achieves this from a single long data record by splitting it into non-overlapping segments.
Given N total samples and a segment length of M, Bartlett's method creates K = ⌊N/M⌋ non-overlapping segments. The periodogram is computed for each segment, then the K periodograms are averaged:
Bartlett's method reveals the fundamental variance-resolution trade-off in PSD estimation. Shorter segments → more segments → lower variance, but poorer frequency resolution. Longer segments → fewer segments → better resolution, but higher variance. You cannot have both simultaneously with a fixed amount of data.
Welch's Method — Overlapping Windowed Segments
Peter Welch (1967) improved Bartlett's method in two key ways: overlapping segments and window functions. By allowing segments to overlap, more periodograms can be computed from the same data, increasing K and further reducing variance without throwing away data at segment edges.
The window function applied to each segment reduces spectral leakage — the bleeding of energy from a strong frequency component into neighboring bins. Without windowing (i.e., the rectangular window), the sharp edges at segment boundaries create large sidelobes in the frequency domain. Applying a smooth window (Hann, Hamming, Blackman) tapers the edges and dramatically reduces leakage.
Bartlett vs. Welch — Key Differences
- Overlapping segments (typically 50%)
- Smooth window function applied
- More segments for same data length
- Greater variance reduction
- Reduced spectral leakage
- Industry standard (pwelch, scipy)
Choosing Welch Parameters in Practice
Three parameters control the Welch estimator's behavior. Choosing them correctly is an engineering judgment that depends on your application:
| Parameter | Effect on Resolution | Effect on Variance | Typical Value |
|---|---|---|---|
| Segment length M | ↑M → finer Δf = f_s/M | ↑M → fewer segments → higher variance | 256–2048 samples |
| Overlap % | No effect on resolution | ↑overlap → more segments → lower variance | 50% (Hann) or 75% |
| Window type | Wider main lobe = coarser resolution | Lower sidelobes = less leakage bias | Hann (general purpose) |
With a Hann window and 50% overlap, consecutive segments share half their data. This overlap percentage is optimal because the Hann window has zero amplitude at the edges — samples near segment boundaries are downweighted and overlap lets them contribute fully in the adjacent segment. Increasing overlap beyond 50–75% provides diminishing returns since successive segments become highly correlated.
Hann window + 50% overlap = the most common Welch configurationImplementation — pwelch and scipy
Welch's method is the default PSD estimation tool in MATLAB (pwelch) and Python (scipy.signal.welch). Both functions implement the same algorithm but use slightly different default conventions:
| Parameter | MATLAB pwelch | scipy.signal.welch |
|---|---|---|
| Default window | Hamming | Hann |
| Default segment length | ⌊N/8⌋ | 256 |
| Default overlap | 50% | 50% |
| Output convention | One-sided by default | One-sided by default |
| Normalization | Power per Hz (density) | Power per Hz (density) |
A key subtlety: both tools output a one-sided PSD by default with units of power/Hz (also called spectral density). If you want total power in a frequency band, integrate the PSD over that band: P_band = ∫ S(f) df ≈ Σ S[k] · Δf.
When comparing PSD plots from different sources, always check: (1) one-sided vs. two-sided, (2) linear vs. dB scale, (3) power/Hz vs. power/sample (normalized). A 3 dB offset often indicates one-sided/two-sided confusion; a frequency axis mismatch indicates normalized vs. physical frequency.
Always verify: one-sided or two-sided? Power/Hz or power/sample?- The periodogram (|DFT|²/N) is asymptotically unbiased but has constant variance ≈ S²(ω) — it is an inconsistent estimator.
- Bartlett's method splits data into K non-overlapping segments and averages their periodograms, reducing variance by K at the cost of frequency resolution.
- Welch's method adds two improvements: overlapping segments (more segments → lower variance) and window functions (reduce spectral leakage).
- The three Welch parameters — segment length M, overlap %, and window — trade off frequency resolution against variance reduction.
- Hann window with 50% overlap is the most common practical configuration; used by default in pwelch and scipy.signal.welch.
- Welch output is one-sided by default; integrate over a band to compute band power.
- There is no free lunch: given fixed total data, you cannot simultaneously improve resolution and reduce variance.