LA 101
M10 · L02
Module 10: Convolution & Filters

Filter Design with Linear Algebra

Filters are matrices. Designing them optimally means asking: which matrix produces the best output? Least squares, the Wiener filter, adaptive algorithms, and beamforming all answer this question using the same elegant linear algebra toolkit.

01 / 12
LA 101
M10 · L02
The FIR Filter

Filter = Matrix-Vector Product

An M-tap FIR filter with coefficients h maps a block of input samples to output via a Toeplitz matrix multiplication: y = Hx. Filter design is then a matrix design problem — choose the best H.

The Design Problem
Given input x and desired output d, find the coefficient vector h that minimizes the difference between Hx and d. This is the core question of optimal filter design.
02 / 12
LA 101
M10 · L02
Optimal Design

Least Squares Filter Design

Collect N input windows as rows of data matrix X (N×M) and the desired samples in vector d (N×1). Minimize the squared error:

Least Squares Cost
J(\mathbf{h})=\|\mathbf{d}-X\mathbf{h}\|^2

This is a quadratic bowl in h — one unique global minimum, found by setting the gradient to zero.

03 / 12
LA 101
M10 · L02
The Solution

The Normal Equations

Setting ∂J/∂h = 0 yields the normal equations — M linear equations whose solution is the optimal filter:

Normal Equations
(X^T X)\,\mathbf{h}^*=X^T\mathbf{d}

XᵀX is the empirical autocorrelation matrix. The solution projects d onto the column space of X — a geometric interpretation of optimality.

04 / 12
LA 101
M10 · L02
Statistical Optimality

The Wiener Filter

When signals are modeled as random processes, minimizing expected MSE yields the Wiener-Hopf equation — the statistical counterpart of the normal equations:

Wiener-Hopf Equation
R\,\mathbf{h}_{\mathrm{opt}}=\mathbf{p}

R is the input autocorrelation matrix; p is the cross-correlation with the desired signal. R is symmetric positive definite Toeplitz — solvable in O(M²) via Levinson-Durbin.

05 / 12
LA 101
M10 · L02
Eigenstructure

Eigenvalues of R Govern Performance

Because R is symmetric positive definite, R = QΛQᵀ. The eigenvalues reveal the signal's power distribution across M directions. The condition number κ(R) = λ_max/λ_min tells the full story:

  • Large κ(R): narrow-band input → ill-conditioned filter design
  • Small κ(R): broad-band input → well-conditioned, stable solution
  • Low-power eigenvectors amplify noise in h_opt
  • Regularization (ridge regression) stabilizes ill-conditioned cases
06 / 12
LA 101
M10 · L02
Adaptive Filtering

LMS: Online Gradient Descent

When statistics change over time, adaptive filters update online using a single sample gradient estimate. LMS is gradient descent with noisy gradients:

LMS Update
\mathbf{h}[n+1]=\mathbf{h}[n]+\mu\,e[n]\,\mathbf{x}[n]

Cost: O(M) per step. No matrix inversions. Convergence speed governed by the condition number κ(R) — ill-conditioned R means slow LMS convergence.

07 / 12
LA 101
M10 · L02
Faster Adaptation

RLS: Exact Least Squares, Recursive

RLS minimizes the full weighted history of errors. It updates the inverse autocorrelation P = (XᵀX)⁻¹ recursively via the matrix inversion lemma — no full matrix inversion at each step.

O(M)
LMS Cost
O(M²)
RLS Cost
M steps
RLS Convergence
08 / 12
LA 101
M10 · L02
Antenna Arrays

Beamforming: Constrained Optimization

An antenna array applies weights w to K antennas. The MVDR beamformer minimizes output power (suppressing noise) while maintaining unit gain in the desired direction — a constrained quadratic program:

MVDR Beamformer
\min_{\mathbf{w}}\;\mathbf{w}^H R_z\mathbf{w}\quad\text{s.t.}\quad\mathbf{w}^H\mathbf{a}(\theta_0)=1
09 / 12
LA 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

10 / 12
LA 101
Key Takeaways
Summary

Key Takeaways

  • FIR filter = Toeplitz matrix; design = least-squares via normal equations (XᵀX)h = Xᵀd
  • Wiener filter: statistical version — Rh = p, where R is autocorrelation, p is cross-correlation
  • Condition number κ(R) determines filter robustness and adaptive convergence speed
  • LMS: O(M) per update, simple gradient step; RLS: O(M²) per update, exact in M steps
  • Beamforming is a constrained quadratic optimization solvable with a single matrix inverse
11 / 12
LA 101
Up Next
Coming Up

M10-L3: PCA and Dimensionality Reduction

We've designed filters by choosing optimal eigenvalues. Now we turn that lens on data itself: the covariance matrix of a dataset has eigenvectors (principal components) that reveal the directions of maximum variance — compressing, denoising, and visualizing high-dimensional data.

12 / 12