DSP 101
M6 · L3
Module 6 — The Fast Fourier Transform
FFT in Practice

Power-of-2 sizes, zero-padding, real-valued optimizations, overlap-add filtering, and the libraries that make it all fast — the engineering decisions behind every working FFT deployment.

1 / 9
DSP 101
M6 · L3
The Golden Rule
Always Use Power-of-2 Sizes

Radix-2 Cooley-Tukey only works when N = 2^m. Libraries are heavily optimized for these sizes. Common choices:

256
low latency
1024
general audio
4096
LTE / 5G OFDM
2 / 9
DSP 101
M6 · L3
Sizing Up
Zero-Padding to the Next Power of 2

Signal block has 700 samples? Append 324 zeros → 1024. Zero-padding interpolates the spectrum (more bins, smoother curve) but does not improve frequency resolution.

Resolution truth
Resolution depends on actual signal length M, not FFT size N. More data = better resolution.
3 / 9
DSP 101
M6 · L3
Free 2× Speedup
The Real-Valued FFT

Real input x[n] produces conjugate-symmetric output: X[N−k] = X*[k]. Only bins 0 … N/2 carry unique information. Pack two real N/2 signals into one complex N-point FFT — halves the work.

Conjugate Symmetry
X[N-k] = X^*[k]
4 / 9
DSP 101
M6 · L3
Filtering Long Signals
Overlap-Add Method
  • Divide x[n] into non-overlapping L-sample blocks
  • Zero-pad each block to N = L + M − 1 (next power of 2)
  • FFT block, multiply by H[k], IFFT back
  • Overlap & add last M−1 samples with next block
  • Result = exact linear convolution with h[n]
  • Cost: O(N log N) per block — not O(N·M)
5 / 9
DSP 101
M6 · L3
The Sibling Method
Overlap-Save

Blocks overlap by M−1 input samples. After each IFFT, the first M−1 output samples (circular convolution artifacts) are discarded. No addition step needed — preferred in hardware.

OLA vs OLS
Identical output when done correctly. OLS discards; OLA adds. Both run at O(N log N).
6 / 9
DSP 101
M6 · L3
Use the Ecosystem
FFT Libraries & Hardware
  • FFTW — auto-tunes for any CPU; gold standard for C/C++
  • Intel oneMKL / IPP — AVX-512 optimized, fastest on Intel
  • NumPy rfft / SciPy — Python; uses FFTPACK or pocketfft
  • cuFFT / rocFFT — hundreds of parallel FFTs on GPU
  • ARM Ne10, Apple Accelerate — NEON / vDSP for mobile
  • DSP ASICs — 1024-point FFT in under 1 microsecond
7 / 9
DSP 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

8 / 9
DSP 101
M6 · L3
Key Takeaways
What You Learned
  • Always use power-of-2 N — zero-pad to the next one
  • Zero-pad = denser bins, not better resolution
  • Real input → use rfft; conjugate symmetry → 2× faster
  • Overlap-Add / Overlap-Save: O(N log N) filtering of long signals
  • FFTW, MKL, cuFFT, Accelerate — use libraries, not hand-rolled code
  • Hardware accelerators: sub-microsecond FFT in comms ASICs
Read full article →
9 / 9
1 / 8 muted