Divide an N-point DFT into two N/2-point DFTs, repeat recursively, and cut the complex multiplications from N² to (N/2)·log₂N — a speedup of up to 100,000× for large signals.
Separate input samples into even-indexed (x[0], x[2], x[4], …) and odd-indexed (x[1], x[3], x[5], …). Each group is an N/2-point DFT. Combine with twiddle factors.
One complex multiply (the twiddle factor W_N^k), two adds, and two outputs. Because W^{k+N/2} = −W^k, we get both X[k] and X[k+N/2] from a single multiply — the efficiency key.
Keep halving: N → N/2 → N/4 → … → 2. At the bottom, a 2-point DFT is just one butterfly with W = 1. The recursion tree has log₂ N levels, each requiring N/2 butterflies.
The speedup grows with N — making the FFT increasingly superior for the large transforms that modern applications demand.
The recursive even-odd splitting reorders input samples by reversing their binary index. After bit-reversal, all log₂ N butterfly stages run in-place — no extra memory needed beyond the N-point input array.
- Radix-2 DIT: classic, split into two halves, (N/2)log₂N multiplies
- Radix-2 DIF: split output bins instead; same count, reversed order
- Radix-4: split into four quarters; ~25% fewer multiplications
- Split-radix: radix-2 for evens, radix-4 for odds — theoretical minimum
- FFTW / Intel IPP: auto-select best algorithm for the hardware
- All are Cooley-Tukey at heart — divide, twiddle, combine
- Split N-point DFT into even/odd halves — two N/2-point DFTs
- Butterfly: 1 multiply + 2 adds → 2 output bins (W^{k+N/2} = −W^k)
- log₂ N stages × N/2 butterflies = O(N log N) total
- N = 1,024: 5,120 vs. 1,048,576 multiplies — 205× faster
- In-place via bit-reversal permutation — O(N) memory
- Split-radix achieves minimum known multiply count