Power-of-2 sizes, zero-padding, real-valued optimizations, overlap-add filtering, and the libraries that make it all fast — the engineering decisions behind every working FFT deployment.
Radix-2 Cooley-Tukey only works when N = 2^m. Libraries are heavily optimized for these sizes. Common choices:
Signal block has 700 samples? Append 324 zeros → 1024. Zero-padding interpolates the spectrum (more bins, smoother curve) but does not improve frequency resolution.
Real input x[n] produces conjugate-symmetric output: X[N−k] = X*[k]. Only bins 0 … N/2 carry unique information. Pack two real N/2 signals into one complex N-point FFT — halves the work.
- Divide x[n] into non-overlapping L-sample blocks
- Zero-pad each block to N = L + M − 1 (next power of 2)
- FFT block, multiply by H[k], IFFT back
- Overlap & add last M−1 samples with next block
- Result = exact linear convolution with h[n]
- Cost: O(N log N) per block — not O(N·M)
Blocks overlap by M−1 input samples. After each IFFT, the first M−1 output samples (circular convolution artifacts) are discarded. No addition step needed — preferred in hardware.
- FFTW — auto-tunes for any CPU; gold standard for C/C++
- Intel oneMKL / IPP — AVX-512 optimized, fastest on Intel
- NumPy rfft / SciPy — Python; uses FFTPACK or pocketfft
- cuFFT / rocFFT — hundreds of parallel FFTs on GPU
- ARM Ne10, Apple Accelerate — NEON / vDSP for mobile
- DSP ASICs — 1024-point FFT in under 1 microsecond
- Always use power-of-2 N — zero-pad to the next one
- Zero-pad = denser bins, not better resolution
- Real input → use rfft; conjugate symmetry → 2× faster
- Overlap-Add / Overlap-Save: O(N log N) filtering of long signals
- FFTW, MKL, cuFFT, Accelerate — use libraries, not hand-rolled code
- Hardware accelerators: sub-microsecond FFT in comms ASICs