Computational Bayesian Methods
The posterior is easy to write down and, in general, impossible to normalise. MCMC sidesteps the integral entirely — it samples from the posterior instead of evaluating it.
The Intractable Denominator
Outside conjugate pairs that integral over θ has no closed form, and in many dimensions no numerical quadrature can reach it. We can compute the numerator for any θ — never the constant that normalises it.
A Chain That Settles on the Posterior
- Build a Markov chain — each step depends only on the current state
- Choose its transitions so its stationary distribution is the posterior itself
- Run it long enough and the states it visits are draws from p(θ∣D)
- Any posterior quantity — mean, credible interval, tail probability — is then an average over those draws
Crucially, only ratios of the posterior are ever needed, so the missing normalising constant cancels and never has to be computed.
Metropolis–Hastings
Propose a move x→x′, then accept it with probability α — otherwise stay put. For a symmetric proposal the q ratio cancels to α = min(1, p(x′)/p(x)). Detailed balance is what makes the posterior stationary.
Proposal Width: Acceptance vs. Mixing
- Too small — nearly every move accepted, but the chain shuffles in place: high autocorrelation, tiny effective sample size
- Too large — nearly every move rejected, so the chain sticks and stalls
- A healthy band between — useful-sized moves accepted often enough to explore
- In high dimensions an acceptance rate near 0.234 is a well-known rule of thumb (Roberts, Gelman & Gilks, 1997)
Gibbs Sampling
When every parameter’s full conditional is known in closed form, draw each coordinate in turn from its own conditional with the others fixed. Every draw is accepted — no rejection step. For the correlated bivariate normal above the conditionals are exactly Gaussian. Gibbs mixes well where you can use it, but the conditionals cannot always be derived.
Trace Plots & Convergence
- Trace plot — a healthy chain is a fuzzy horizontal band; a stuck one shows long flat runs or slow drift
- Burn-in — discard the early draws taken before the chain reached its stationary region
- Autocorrelation & ESS — correlated draws carry less information; the effective sample size is how many independent draws they are worth
- R-hat — run several chains from different starts; when they overlap (R̂ ≈ 1) they have converged
Watch a Chain Mix
Move the proposal width and watch acceptance trade off against mixing — small steps crawl, large steps stick, and a middle band explores.
What you learned
- The posterior’s normalising integral is usually intractable — MCMC samples the posterior instead of evaluating it
- Metropolis–Hastings proposes, then accepts with α = min(1, …); detailed balance makes the posterior stationary
- Proposal width trades acceptance against mixing; ~0.234 is the high-dimensional rule of thumb
- Gibbs draws each coordinate from its exact full conditional and accepts every draw
- Diagnose with trace plots, burn-in, autocorrelation/ESS and R-hat across chains