Reading
Stories Mode

Sample Rate Conversion

~13 min read Lesson 3 of Module 11

Converting Between Any Two Rates

The previous two lessons treated upsampling and downsampling as separate, integer-factor operations. Real systems, however, often need to convert between sample rates that are not integer multiples of each other. A professional audio workstation records at 48 kHz but must deliver files at 44.1 kHz for CD mastering. A software-defined radio running at 30.72 MHz must interface with a baseband processor clocked at 1.92 MHz. In both cases, the conversion ratio is a rational number L/M, not an integer.

Sample rate conversion (SRC) by L/M is achieved by cascading an interpolator (upsample by L) with a decimator (downsample by M), sharing a single combined filter between them. This elegant structure avoids operating at the intermediate high rate and produces an output sample rate of exactly (L/M) × fs from an input at fs.

Core Idea

Upsample by L, then filter with a single lowpass whose cutoff is min(π/L, π/M), then downsample by M. The filter simultaneously prevents aliasing (for the decimation step) and removes imaging (from the upsampling step).

fout = (L / M) × fin

The Cascade Structure

In principle, an SRC system could use separate anti-imaging and anti-aliasing filters: one after the upsampler and one before the downsampler. But since both are lowpass filters, they can be merged into a single unit. The combined filter specification is:

Combined SRC Filter
H\!\left(e^{j\omega}\right) = \begin{cases} L & |\omega| \le \min\!\left(\dfrac{\pi}{L},\, \dfrac{\pi}{M}\right) \\ 0 & \text{otherwise} \end{cases}
The SRC filter cutoff is the more restrictive of the anti-imaging cutoff (π/L) and the anti-aliasing cutoff (π/M), evaluated at the intermediate high rate Lfs. The gain in the passband is L to compensate for zero insertion.

The system diagram is: x[n] → upsample by L → lowpass filter H → downsample by M → y[m]. At the intermediate rate Lfs, the upsampled signal passes through H before every M-th sample is selected for the output. The result is a sequence at the target rate (L/M)fs.

A Canonical Example: 44.1 kHz ↔ 48 kHz

The conversion between 44.1 kHz and 48 kHz is ubiquitous in professional audio — it arises whenever content created for digital streaming (48 kHz) must be distributed on CD (44.1 kHz), or vice versa. The ratio is:

44.1 ↔ 48 kHz Ratio
\frac{f_{\text{out}}}{f_{\text{in}}} = \frac{48{,}000}{44{,}100} = \frac{160}{147}
To convert from 44.1 kHz to 48 kHz: upsample by L = 160, filter, then downsample by M = 147. The intermediate rate is 160 × 44.1 kHz = 7.056 MHz. To go the other direction, swap L and M.
160
Upsample factor L (44.1→48 kHz)
147
Downsample factor M (44.1→48 kHz)
7.056
Intermediate rate (MHz)

Running the filter at 7.056 MHz naively would be prohibitively expensive. The polyphase structure resolves this: it avoids materializing the intermediate rate entirely, reducing computation to a fraction of the naïve cost.

Designing the Combined Filter

The combined SRC filter must satisfy two simultaneous constraints. When L > M (upsampling dominates), the anti-imaging cutoff π/L is tighter than π/M and governs the design. When M > L (downsampling dominates), the anti-aliasing cutoff π/M is the binding constraint. In both cases the cutoff is set to min(π/L, π/M) and the passband gain is L.

Filter Design Checklist

To design the SRC lowpass filter: (1) set cutoff to π/max(L, M) in radians per sample at the intermediate rate Lfs; (2) set passband gain to L; (3) choose filter order to meet stopband attenuation — typically ≥ 80 dB for audio. A Parks-McClellan equiripple FIR is the standard choice.

The transition bandwidth of the filter is determined by the worst-case band-edge: a signal that occupies the full bandwidth of the input (up to π at fs) must be kept below π/max(L, M) at the output. For audio, a usable transition band exists in [20 kHz, ½fmin], which comfortably accommodates even sharp FIR designs without audible ringing.

Efficient Implementation: Polyphase SRC

Materializing the intermediate rate Lfs and running a long FIR filter there would be impractical. The polyphase approach avoids this by recognizing that the upsample-filter-downsample chain can be reorganized so that all filtering happens at either fs or (L/M)fs — whichever is lower.

The combined FIR filter of length N is decomposed into L polyphase branches. Only every M-th output of the filter is actually needed (due to the subsequent downsampling). By selecting which branch to evaluate and in what order, the polyphase SRC computes exactly those outputs — and no others. The result is an algorithm that processes one output sample by evaluating a single polyphase branch of N/L taps, advancing through the L branches cyclically with a stride of M.

Polyphase SRC Output
y[m] = \sum_{k=0}^{N/L - 1} E_{k(m)}[k]\, x[n(m) - k], \quad k(m) = (mM \bmod L)
Each output sample y[m] is computed from polyphase branch k(m) = (mM mod L), evaluated at input index n(m) = ⌊mM/L⌋. The algorithm steps through branches with stride M, wrapping modulo L, and advances the input pointer whenever it crosses an L boundary.

The computational cost of polyphase SRC is approximately N/L multiplications per output sample, independent of the intermediate rate. For the 44.1↔48 kHz converter with L = 160 and a 320-tap filter, this is just 2 multiplications per output sample — far less than the naïve 320.

Arbitrary and Asynchronous Rate Conversion

Rational SRC with fixed L and M handles any conversion whose ratio can be expressed as a fraction with manageable numerator and denominator. But some scenarios demand asynchronous SRC: the input and output clocks are independent and the ratio drifts over time (e.g., two devices sharing a network with timing jitter). In this case, a fixed L/M cannot be precomputed.

Asynchronous SRC uses a continuously varying fractional delay. The idea is to treat the input sequence as defining a bandlimited continuous-time signal via the sinc interpolation formula, then resample it at the desired output time instants. In practice this is implemented by maintaining a high-resolution time-phase accumulator and evaluating the appropriate polyphase branch (or an interpolation between two adjacent branches) at each output moment.

Professional audio interfaces, network-synchronized audio systems (such as AES67 and Dante), and SDR front-ends all rely on asynchronous SRC to handle real-world clock drift without audible or perceptible artifacts.

Practical Applications

Professional audio: Every Digital Audio Workstation includes a high-quality SRC engine. Converting between 44.1 kHz, 48 kHz, 88.2 kHz, and 96 kHz is routine, and the polyphase SRC filter must achieve ≥ 120 dB stopband attenuation to avoid audible aliasing or imaging in high-resolution audio.

Software-defined radio: An SDR receiver digitizes a wide RF band at a high rate (e.g., 61.44 MHz) and must deliver a narrowband channel at a much lower rate (e.g., 200 kHz for a GSM channel). A cascaded SRC chain — often using successive integer decimations followed by a final rational SRC stage — accomplishes this while suppressing adjacent-channel interference at each step.

Video and multimedia: Video codecs operate at frame rates such as 23.976, 25, 29.97, 30, and 60 fps. Transcoding between these rates requires temporal SRC; the same polyphase framework applies in the time domain (frames) rather than the sample domain.

Next lesson: Filter Banks — partitioning a wideband signal into multiple subbands using analysis and synthesis filter banks, the foundation of audio codecs and OFDM systems.

Key Takeaways
  • Sample rate conversion by rational factor L/M is achieved by upsampling by L, filtering with a combined lowpass at cutoff min(π/L, π/M) and gain L, then downsampling by M.
  • The single combined filter replaces both the anti-imaging filter (from upsampling) and the anti-aliasing filter (from downsampling), using the tighter of the two cutoffs.
  • The canonical 44.1 ↔ 48 kHz conversion uses L = 160 and M = 147, yielding an intermediate rate of 7.056 MHz that is never physically realized when using polyphase implementation.
  • Polyphase SRC decomposes the filter into L branches and evaluates only the branches needed for each output sample, reducing cost to N/L multiplications per output regardless of the intermediate rate.
  • Asynchronous SRC handles clock-independent sources by continuously updating a fractional phase accumulator and evaluating the appropriate polyphase branch at each output instant.
  • SRC is fundamental in professional audio, software-defined radio, and multimedia transcoding wherever signals must cross sample-rate boundaries.
Previous Interpolation (Upsampling) Module Overview Next Lesson Filter Banks