DSP 101
M12 · L01
Module 12: Capstone & Real-World DSP

DSP on Embedded Systems

Theory meets silicon. How do filters, FFTs, and correlators actually run on a chip with milliwatts of power and microseconds to spare?

01 / 11
DSP 101
M12 · L01
The Reality

Three Hard Constraints

Embedded DSP is a constant three-way trade-off. Unlike desktop software, you cannot buy your way out of any of these.

mW
Power budget — battery life depends on it
KB
Memory — kilobytes, not gigabytes
µs
Deadline — miss it and the signal corrupts
02 / 11
DSP 101
M12 · L01
Arithmetic

Fixed-Point vs Floating-Point

Floating-point is easy to code but expensive in silicon. Fixed-point (Q15, Q31) runs on integer hardware — 10× less power, same algorithm.

Q-Format
x = k \cdot 2^{-N}
03 / 11
DSP 101
M12 · L01
The Danger

Overflow is Catastrophic

  • Fixed-point has a finite range — exceed it and the value wraps
  • Wrapping in audio: a loud click or total distortion
  • Solution: saturating arithmetic — clamp to max, not wrap
  • Solution: careful scaling — leave headroom before multiplying
  • 64-bit accumulator absorbs hundreds of products safely
04 / 11
DSP 101
M12 · L01
Hardware

Processors Built for DSP

Not all CPUs are equal. DSP cores add features that make signal loops dramatically faster.

  • Single-cycle MAC: y += a × b in one clock
  • SIMD: two (or eight) MACs per cycle
  • Circular buffer addressing: delay lines at zero cost
  • Bit-reversed addressing: FFT butterfly reordering in hardware
05 / 11
DSP 101
M12 · L01
Real Devices

TI C6000 & ARM Cortex-M

TI C6748
3,000 MMACS · VLIW · <1 W · audio, medical, modem
ARM Cortex-M55
Helium SIMD · 8 MACs/cycle · mA supply · wearables, IoT
06 / 11
DSP 101
M12 · L01
Memory Hierarchy

Keep Data Close to the Core

Off-chip DRAM is 10–100× slower than on-chip SRAM. Coefficients and sample buffers must live in TCM or L1 cache to meet real-time deadlines.

Example Budget
256-tap Q15 FIR coefficients = 512 B · 1024-pt complex FFT buffer = 8 KB · both fit in Cortex-M7 TCM
07 / 11
DSP 101
M12 · L01
Real-Time

Predictably Fast

Real-time is not about being fast on average — it means never missing a deadline. Even a single overrun corrupts the signal.

CPU Load
\text{Load} = \frac{\text{cycles / frame}}{f_{\mathrm{clk}} \times T_{\mathrm{frame}}}
08 / 11
DSP 101
Knowledge Check

Check whatstuck

Four questions on embedded DSP — fixed-point arithmetic, overflow, DSP-core features, and the memory hierarchy.

Question 1 of 0
Score 0/0

09 / 11
DSP 101
M12 · L01
Optimization

Code That Runs Fast

  • Use CMSIS-DSP / DSPLIB: hand-tuned SIMD assembly — don't rewrite it
  • SIMD intrinsics: two Q15 MACs per cycle via __SMLAD()
  • Eliminate branches in hot loops — pipeline stalls kill throughput
  • DMA double-buffering: compute while DMA fills the next block
  • Target <70% CPU load — leave headroom for interrupts
10 / 11
DSP 101
Up Next
Coming Up

Audio Effects Pipeline

Put it all together: equalizer, reverb, compression, and delay — a complete multi-stage audio processor from input to output.

11 / 11