Converting Between Any Two Rates
The previous two lessons treated upsampling and downsampling as separate, integer-factor operations. Real systems, however, often need to convert between sample rates that are not integer multiples of each other. A professional audio workstation records at 48 kHz but must deliver files at 44.1 kHz for CD mastering. A software-defined radio running at 30.72 MHz must interface with a baseband processor clocked at 1.92 MHz. In both cases, the conversion ratio is a rational number L/M, not an integer.
Sample rate conversion (SRC) by L/M is achieved by cascading an interpolator (upsample by L) with a decimator (downsample by M), sharing a single combined filter between them. This elegant structure avoids operating at the intermediate high rate and produces an output sample rate of exactly (L/M) × fs from an input at fs.
Upsample by L, then filter with a single lowpass whose cutoff is min(π/L, π/M), then downsample by M. The filter simultaneously prevents aliasing (for the decimation step) and removes imaging (from the upsampling step).
fout = (L / M) × finThe Cascade Structure
In principle, an SRC system could use separate anti-imaging and anti-aliasing filters: one after the upsampler and one before the downsampler. But since both are lowpass filters, they can be merged into a single unit. The combined filter specification is:
The system diagram is: x[n] → upsample by L → lowpass filter H → downsample by M → y[m]. At the intermediate rate Lfs, the upsampled signal passes through H before every M-th sample is selected for the output. The result is a sequence at the target rate (L/M)fs.
A Canonical Example: 44.1 kHz ↔ 48 kHz
The conversion between 44.1 kHz and 48 kHz is ubiquitous in professional audio — it arises whenever content created for digital streaming (48 kHz) must be distributed on CD (44.1 kHz), or vice versa. The ratio is:
Running the filter at 7.056 MHz naively would be prohibitively expensive. The polyphase structure resolves this: it avoids materializing the intermediate rate entirely, reducing computation to a fraction of the naïve cost.
Designing the Combined Filter
The combined SRC filter must satisfy two simultaneous constraints. When L > M (upsampling dominates), the anti-imaging cutoff π/L is tighter than π/M and governs the design. When M > L (downsampling dominates), the anti-aliasing cutoff π/M is the binding constraint. In both cases the cutoff is set to min(π/L, π/M) and the passband gain is L.
To design the SRC lowpass filter: (1) set cutoff to π/max(L, M) in radians per sample at the intermediate rate Lfs; (2) set passband gain to L; (3) choose filter order to meet stopband attenuation — typically ≥ 80 dB for audio. A Parks-McClellan equiripple FIR is the standard choice.
The transition bandwidth of the filter is determined by the worst-case band-edge: a signal that occupies the full bandwidth of the input (up to π at fs) must be kept below π/max(L, M) at the output. For audio, a usable transition band exists in [20 kHz, ½fmin], which comfortably accommodates even sharp FIR designs without audible ringing.
Efficient Implementation: Polyphase SRC
Materializing the intermediate rate Lfs and running a long FIR filter there would be impractical. The polyphase approach avoids this by recognizing that the upsample-filter-downsample chain can be reorganized so that all filtering happens at either fs or (L/M)fs — whichever is lower.
The combined FIR filter of length N is decomposed into L polyphase branches. Only every M-th output of the filter is actually needed (due to the subsequent downsampling). By selecting which branch to evaluate and in what order, the polyphase SRC computes exactly those outputs — and no others. The result is an algorithm that processes one output sample by evaluating a single polyphase branch of N/L taps, advancing through the L branches cyclically with a stride of M.
The computational cost of polyphase SRC is approximately N/L multiplications per output sample, independent of the intermediate rate. For the 44.1↔48 kHz converter with L = 160 and a 320-tap filter, this is just 2 multiplications per output sample — far less than the naïve 320.
Arbitrary and Asynchronous Rate Conversion
Rational SRC with fixed L and M handles any conversion whose ratio can be expressed as a fraction with manageable numerator and denominator. But some scenarios demand asynchronous SRC: the input and output clocks are independent and the ratio drifts over time (e.g., two devices sharing a network with timing jitter). In this case, a fixed L/M cannot be precomputed.
Asynchronous SRC uses a continuously varying fractional delay. The idea is to treat the input sequence as defining a bandlimited continuous-time signal via the sinc interpolation formula, then resample it at the desired output time instants. In practice this is implemented by maintaining a high-resolution time-phase accumulator and evaluating the appropriate polyphase branch (or an interpolation between two adjacent branches) at each output moment.
Professional audio interfaces, network-synchronized audio systems (such as AES67 and Dante), and SDR front-ends all rely on asynchronous SRC to handle real-world clock drift without audible or perceptible artifacts.
Practical Applications
Professional audio: Every Digital Audio Workstation includes a high-quality SRC engine. Converting between 44.1 kHz, 48 kHz, 88.2 kHz, and 96 kHz is routine, and the polyphase SRC filter must achieve ≥ 120 dB stopband attenuation to avoid audible aliasing or imaging in high-resolution audio.
Software-defined radio: An SDR receiver digitizes a wide RF band at a high rate (e.g., 61.44 MHz) and must deliver a narrowband channel at a much lower rate (e.g., 200 kHz for a GSM channel). A cascaded SRC chain — often using successive integer decimations followed by a final rational SRC stage — accomplishes this while suppressing adjacent-channel interference at each step.
Video and multimedia: Video codecs operate at frame rates such as 23.976, 25, 29.97, 30, and 60 fps. Transcoding between these rates requires temporal SRC; the same polyphase framework applies in the time domain (frames) rather than the sample domain.
- Sample rate conversion by rational factor L/M is achieved by upsampling by L, filtering with a combined lowpass at cutoff min(π/L, π/M) and gain L, then downsampling by M.
- The single combined filter replaces both the anti-imaging filter (from upsampling) and the anti-aliasing filter (from downsampling), using the tighter of the two cutoffs.
- The canonical 44.1 ↔ 48 kHz conversion uses L = 160 and M = 147, yielding an intermediate rate of 7.056 MHz that is never physically realized when using polyphase implementation.
- Polyphase SRC decomposes the filter into L branches and evaluates only the branches needed for each output sample, reducing cost to N/L multiplications per output regardless of the intermediate rate.
- Asynchronous SRC handles clock-independent sources by continuously updating a fractional phase accumulator and evaluating the appropriate polyphase branch at each output instant.
- SRC is fundamental in professional audio, software-defined radio, and multimedia transcoding wherever signals must cross sample-rate boundaries.