From Algorithm to Silicon
Everything covered in this course — filters, transforms, correlation, multirate processing — ultimately runs on hardware. In real-world products, that hardware is typically a small, power-constrained embedded processor rather than a desktop workstation. Understanding how DSP algorithms map onto embedded hardware is the bridge between theory and the millions of devices that process signals around us every day: headphones, smartphones, medical monitors, motor controllers, and satellite receivers.
Embedded DSP comes with hard constraints that software engineers rarely face on general-purpose computers: limited memory (often kilobytes, not gigabytes), strict power budgets, and unforgiving real-time deadlines. Missing a sample period means corrupted audio, a dropped communication packet, or a missed motor control update. Navigating these constraints requires understanding arithmetic formats, processor architectures, and code optimization techniques.
Embedded DSP is a three-way trade-off between computation speed, power consumption, and numerical precision. Every design decision tilts the balance. Choosing fixed-point over floating-point, for example, cuts power by 10× but demands careful scaling to avoid overflow and quantization noise.
Fixed-Point vs. Floating-Point Arithmetic
The most fundamental choice in embedded DSP is how numbers are represented. Floating-point numbers (IEEE 754, 32-bit single precision) encode a wide dynamic range automatically — the decimal point “floats” to wherever it is needed. This makes algorithm development straightforward: you rarely worry about overflow or scaling. The penalty is hardware cost: a floating-point unit (FPU) consumes significant chip area and power.
Fixed-point numbers represent values as scaled integers. A Q15 format, for example, stores a 16-bit integer where the implied decimal point sits after the sign bit and before 15 fractional bits, giving a range of −1 to just under +1 with a resolution of 2−15 ≈ 30.5 µV if the full scale is 1 V. Fixed-point arithmetic requires no FPU and executes on simple integer hardware, consuming far less power and silicon area.
The principal risk of fixed-point arithmetic is overflow: if a multiplication or accumulation produces a result outside the representable range, the number wraps around catastrophically. Careful scaling — choosing where to place the binary point and how much headroom to leave — is the central skill of fixed-point DSP programming. Saturating arithmetic (clamping to the maximum representable value instead of wrapping) is a hardware feature on most DSP cores that mitigates overflow gracefully.
DSP Processor Architectures
General-purpose processors (ARM Cortex-A, x86) are optimized for code branching and memory access patterns typical of operating systems and applications. DSP processors are instead optimized for the repetitive, data-intensive loops at the heart of signal processing: multiply-accumulate operations, circular buffering, and bit-reversed addressing for FFTs.
Texas Instruments C6000 family (e.g., C6748, C66x) are high-performance VLIW (Very Long Instruction Word) DSPs with multiple execution units operating in parallel. A single C6748 core can sustain 3,000 million multiply-accumulate operations per second (3,000 MMACS) while consuming under 1 W. They are ubiquitous in audio processing, medical imaging, and baseband modems.
ARM Cortex-M microcontrollers (M4, M7, M33, M55) bring DSP capability to ultra-low-power microcontrollers. The Cortex-M4 added a single-precision FPU and SIMD (single-instruction, multiple-data) instructions — executing two 16-bit multiplications with a single instruction. The Cortex-M55 adds the Helium vector extension, enabling eight simultaneous 16-bit MAC operations. These devices run from milliamps of supply current, making them ideal for battery-powered wearables and IoT sensors.
Every DSP processor includes a dedicated Multiply-Accumulate (MAC) unit that computes y += a × b in a single clock cycle. Convolution, filtering, and correlation all reduce to repeated MAC operations. A 32-bit MAC with a 64-bit accumulator (common in 16-bit fixed-point DSPs) prevents overflow during the accumulation of hundreds of products before the final result is scaled and stored.
Memory and Computation Constraints
Memory in embedded DSP is scarce, expensive (in power and die area), and architecturally important. DSP processors typically provide separate instruction and data memories (Harvard architecture) so that the processor can fetch the next instruction while simultaneously loading a data value — avoiding the von Neumann bottleneck.
Filter coefficient tables and circular sample buffers must fit in on-chip SRAM or tightly coupled memory (TCM) to avoid the 10–100× latency penalty of off-chip DRAM access. For a 256-tap FIR filter at Q15, the coefficient table consumes 512 bytes — easily on-chip. A 1024-point FFT working buffer in 32-bit complex format requires 8 KB — still on-chip for most Cortex-M7 parts but worth budgeting carefully.
Circular (ring) buffers are the standard data structure for implementing delay lines in filters and echo cancellers. A pointer wraps modulo the buffer length, avoiding the cost of shifting the entire array with each new sample. Most DSP processor architectures support hardware-modulo addressing that implements circular buffering with zero overhead.
Real-Time Processing Requirements
A real-time DSP system must process every input sample (or block of samples) within a strict deadline. For audio at 48 kHz, the processor has exactly 20.83 µs per sample. For a 5G baseband processor at 100 MHz symbol rate, the deadline is 10 ns. Missing the deadline causes a buffer overrun: the new sample overwrites one that hasn't been processed yet, producing an audible click, a dropped symbol, or a control fault.
Block processing amortizes interrupt overhead by collecting samples into frames (e.g., 64 or 256 samples) and processing the entire frame at once. The latency (delay from input to output) increases to one frame duration — acceptable for most audio and communications applications — but throughput efficiency improves dramatically. The FFT is inherently a block algorithm and fits naturally into this model.
Real-time is not about being fast — it is about being predictably fast. A system that processes most frames in 5 µs but occasionally takes 25 µs will overrun. Interrupt latency, cache misses, DMA transfers, and branch mispredictions all contribute to worst-case execution time (WCET). Embedded DSP code is written to minimize and bound these variations, not just the average case.
Code Optimization for DSP
Writing DSP code that meets real-time deadlines on power-constrained hardware requires a disciplined optimization process. The most impactful techniques, roughly in order of importance:
Use vendor-optimized libraries. ARM CMSIS-DSP and TI DSPLIB provide hand-tuned assembly implementations of FIR filters, FFTs, matrix operations, and more. A library FFT is typically 5–10× faster than a straightforward C implementation because it exploits SIMD, unrolling, and cache-friendly memory access patterns. Never rewrite what a library already does better.
Exploit SIMD and intrinsics. Modern DSP cores can process multiple data lanes in parallel. ARM Cortex-M4 SIMD processes two Q15 MACs per cycle using the SMLAD instruction. Writing this in C via compiler intrinsics (e.g., __SMLAD()) avoids assembly while still generating optimal code. A hand-written inner loop using SIMD can double or quadruple throughput compared to scalar code.
Avoid data-dependent branching in hot loops. Conditional jumps in inner loops cause pipeline stalls. Saturating arithmetic, predicated instructions, and loop unrolling eliminate branches from the critical path. The compiler flag -funroll-loops helps, but manual unrolling of fixed-iteration loops (e.g., a 32-tap FIR) is often more effective.
Manage memory hierarchy aggressively. Coefficients and data buffers should live in the fastest available memory (L1 SRAM or TCM). DMA (direct memory access) transfers from slower external flash or SDRAM should run in the background, ping-ponging between two buffers so the processor is always computing on a full buffer while DMA fills the next one.
- Fixed-point arithmetic (Q15, Q31) trades dynamic range flexibility for dramatically lower power and hardware cost; careful scaling and saturating arithmetic prevent overflow.
- Dedicated DSP processors (TI C6000) and DSP-enhanced microcontrollers (ARM Cortex-M4/M7/M55) provide single-cycle MAC units, SIMD instructions, and circular buffer addressing that are essential for real-time performance.
- Harvard memory architecture and on-chip SRAM/TCM eliminate the memory bottleneck; coefficients and working buffers must be placed in the fastest available memory tier.
- Real-time processing demands predictable worst-case execution time, not just fast average performance; block processing, DMA double-buffering, and branch elimination are the key techniques.
- Vendor-optimized DSP libraries (CMSIS-DSP, DSPLIB) deliver hand-tuned SIMD performance; use them as a starting point before writing custom assembly.
- CPU load should be kept well below 100% to leave headroom for interrupts, OS overhead, and future features without triggering buffer overruns.