DSP on Embedded Systems
Theory meets silicon. How do filters, FFTs, and correlators actually run on a chip with milliwatts of power and microseconds to spare?
Three Hard Constraints
Embedded DSP is a constant three-way trade-off. Unlike desktop software, you cannot buy your way out of any of these.
Fixed-Point vs Floating-Point
Floating-point is easy to code but expensive in silicon. Fixed-point (Q15, Q31) runs on integer hardware — 10× less power, same algorithm.
Overflow is Catastrophic
- Fixed-point has a finite range — exceed it and the value wraps
- Wrapping in audio: a loud click or total distortion
- Solution: saturating arithmetic — clamp to max, not wrap
- Solution: careful scaling — leave headroom before multiplying
- 64-bit accumulator absorbs hundreds of products safely
Processors Built for DSP
Not all CPUs are equal. DSP cores add features that make signal loops dramatically faster.
- Single-cycle MAC: y += a × b in one clock
- SIMD: two (or eight) MACs per cycle
- Circular buffer addressing: delay lines at zero cost
- Bit-reversed addressing: FFT butterfly reordering in hardware
TI C6000 & ARM Cortex-M
Keep Data Close to the Core
Off-chip DRAM is 10–100× slower than on-chip SRAM. Coefficients and sample buffers must live in TCM or L1 cache to meet real-time deadlines.
Predictably Fast
Real-time is not about being fast on average — it means never missing a deadline. Even a single overrun corrupts the signal.
Code That Runs Fast
- Use CMSIS-DSP / DSPLIB: hand-tuned SIMD assembly — don't rewrite it
- SIMD intrinsics: two Q15 MACs per cycle via
__SMLAD() - Eliminate branches in hot loops — pipeline stalls kill throughput
- DMA double-buffering: compute while DMA fills the next block
- Target <70% CPU load — leave headroom for interrupts
Audio Effects Pipeline
Put it all together: equalizer, reverb, compression, and delay — a complete multi-stage audio processor from input to output.