Filter Design with Linear Algebra
Filters are matrices. Designing them optimally means asking: which matrix produces the best output? Least squares, the Wiener filter, adaptive algorithms, and beamforming all answer this question using the same elegant linear algebra toolkit.
Filter = Matrix-Vector Product
An M-tap FIR filter with coefficients h maps a block of input samples to output via a Toeplitz matrix multiplication: y = Hx. Filter design is then a matrix design problem — choose the best H.
Least Squares Filter Design
Collect N input windows as rows of data matrix X (N×M) and the desired samples in vector d (N×1). Minimize the squared error:
This is a quadratic bowl in h — one unique global minimum, found by setting the gradient to zero.
The Normal Equations
Setting ∂J/∂h = 0 yields the normal equations — M linear equations whose solution is the optimal filter:
XᵀX is the empirical autocorrelation matrix. The solution projects d onto the column space of X — a geometric interpretation of optimality.
The Wiener Filter
When signals are modeled as random processes, minimizing expected MSE yields the Wiener-Hopf equation — the statistical counterpart of the normal equations:
R is the input autocorrelation matrix; p is the cross-correlation with the desired signal. R is symmetric positive definite Toeplitz — solvable in O(M²) via Levinson-Durbin.
Eigenvalues of R Govern Performance
Because R is symmetric positive definite, R = QΛQᵀ. The eigenvalues reveal the signal's power distribution across M directions. The condition number κ(R) = λ_max/λ_min tells the full story:
- Large κ(R): narrow-band input → ill-conditioned filter design
- Small κ(R): broad-band input → well-conditioned, stable solution
- Low-power eigenvectors amplify noise in h_opt
- Regularization (ridge regression) stabilizes ill-conditioned cases
LMS: Online Gradient Descent
When statistics change over time, adaptive filters update online using a single sample gradient estimate. LMS is gradient descent with noisy gradients:
Cost: O(M) per step. No matrix inversions. Convergence speed governed by the condition number κ(R) — ill-conditioned R means slow LMS convergence.
RLS: Exact Least Squares, Recursive
RLS minimizes the full weighted history of errors. It updates the inverse autocorrelation P = (XᵀX)⁻¹ recursively via the matrix inversion lemma — no full matrix inversion at each step.
Beamforming: Constrained Optimization
An antenna array applies weights w to K antennas. The MVDR beamformer minimizes output power (suppressing noise) while maintaining unit gain in the desired direction — a constrained quadratic program:
Key Takeaways
- FIR filter = Toeplitz matrix; design = least-squares via normal equations (XᵀX)h = Xᵀd
- Wiener filter: statistical version — Rh = p, where R is autocorrelation, p is cross-correlation
- Condition number κ(R) determines filter robustness and adaptive convergence speed
- LMS: O(M) per update, simple gradient step; RLS: O(M²) per update, exact in M steps
- Beamforming is a constrained quadratic optimization solvable with a single matrix inverse
M10-L3: PCA and Dimensionality Reduction
We've designed filters by choosing optimal eigenvalues. Now we turn that lens on data itself: the covariance matrix of a dataset has eigenvectors (principal components) that reveal the directions of maximum variance — compressing, denoising, and visualizing high-dimensional data.