Reading
Stories Mode

Non-Parametric Regression

~30 min read Lesson 3 of 3 in Module 8

Regression Without a Formula

Module 7 always began by assuming a form for the relationship — a straight line, a polynomial, a logistic curve — and then estimating its parameters. Non-parametric regression drops that assumption. Instead of forcing the data into a chosen shape, it lets the data draw its own curve, using only a smoothness constraint. It is what you reach for when you genuinely do not know the functional form, or when the relationship is too wiggly for any tidy equation.

Kernel Density Estimation

The idea starts one step before regression, with estimating a distribution. Kernel density estimation (KDE), taught in Module 2, Lesson 1, is the smooth cousin of the histogram — so we recall its formula here rather than re-deriving it:

Kernel Density Estimate
\hat{f}(x) = \dfrac{1}{nh}\sum_{i=1}^{n} K\!\left(\dfrac{x - x_i}{h}\right)
Place a small bump (the kernel K) over every data point and add them up. The bandwidth h sets the width of each bump — the single knob that controls how smooth the estimate is. This same “local averaging” idea, applied to y instead of to counts, is non-parametric regression.

LOESS: Local Regression

LOESS (or LOWESS, locally weighted scatterplot smoothing) fits the curve one point at a time. To find the fitted value at a location x, it runs a little weighted least-squares regression using only nearby points, weighting each by how close it is:

Local Weighted Fit
\min_{a,b}\; \sum_{i=1}^{n} w_i(x)\,\big(y_i - a - b\,x_i\big)^2
The weights wi(x) fall off with distance from x (a kernel again), so far-away points barely count. Slide x across the range, refit at each step, and the fitted values trace a smooth curve. The span — the fraction of points counted as “near” — plays the role of the bandwidth.

Splines: Piecewise Polynomials

A spline takes a different route to the same freedom: it splits the x-axis at a set of knots and fits a low-degree polynomial (usually cubic) between each pair, stitched together so the joins are smooth — matching value and slope at every knot. More knots mean more flexibility. A smoothing spline automates the choice by placing a knot at every point and adding a roughness penalty (a close relative of the Ridge penalty from Module 7) that trades fit against wiggliness. Splines are the workhorse behind the flexible terms in generalized additive models.

Gaussian Process Regression

Gaussian process (GP) regression is the Bayesian member of this family — a natural bridge to the next module. Rather than fitting one curve, a GP puts a prior over functions and returns a posterior distribution over all curves consistent with the data. Its output is therefore not just a fitted line but a full uncertainty band that widens where data is sparse and tightens where data is dense — honest error bars for a shape you never had to specify.

The One Knob: Bandwidth and the Bias-Variance Trade-off

Every method here has a single smoothness dial — bandwidth, span, number of knots, or GP length-scale — and it is the bias-variance trade-off made tangible. Too smooth (wide bandwidth) and the curve underfits, missing real structure: high bias. Too rough (narrow bandwidth) and it chases noise: high variance. Choosing that dial — usually by cross-validation — is the whole art of non-parametric regression. Use these methods when flexibility matters more than a clean formula; use parametric regression when you want interpretable coefficients and have reason to trust the form.

Key Takeaways
  • Non-parametric regression assumes no functional form — the data draws its own curve under a smoothness constraint.
  • Kernel density estimation (Module 2, Lesson 1) is the smooth histogram; the same local-averaging idea, applied to y, gives non-parametric regression.
  • LOESS fits a weighted local regression at each point; its span controls smoothness.
  • Splines stitch low-degree polynomials at knots; smoothing splines add a roughness penalty akin to Ridge.
  • Gaussian process regression is the Bayesian version — a posterior over functions with honest, data-dependent uncertainty bands.
  • All share one smoothness knob, tuned by cross-validation — it is the bias-variance trade-off made concrete.
Previous Non-Parametric Tests Module Overview Next — Module 9 The Bayesian Framework