STAT 101
M9 · Lab
Hands-on lab
Sample a posterior you cannot solve

Run a Metropolis-Hastings sampler on a known target in the sandbox, and predict — before you read it off the screen — the mean, variance and covariance the chain converges to, the exact standard deviation Gibbs draws from, and why the autocorrelation starts at one.

Posterior ∝ likelihood × prior
p(\theta \mid D) \propto p(D \mid \theta)\, p(\theta)

A posterior is the prior times the likelihood, renormalised: p(θ | data) ∝ p(data | θ) · p(θ). When the prior is conjugate to the likelihood — a Beta(a, b) prior with a binomial count of s successes in n trials gives a Beta(a + s, b + n − s) posterior, mean (a + s)/(a + b + n) — the answer is closed-form. Most posteriors are not conjugate and have no closed form, so you sample them: build a Markov chain whose stationary distribution IS the posterior, and treat the draws as if they came from it. The only way to trust the draws is to run the sampler on a target whose truth you already know — here a correlated normal with mean (0, 0), variances 1, covariance ρ — and watch them converge to it.

1 / 9
STAT 101
M9 · Lab
Set it up
One chain

Open the sandbox. It opens on exactly this state — a Metropolis-Hastings chain on the correlated normal — so a Reset gets you here too. Each step names the one knob to change; leave everything else alone, and let the chain run out to tens of thousands of samples before you read a moment.

Set these values
Target -> Correlated bivariate normal Sampler -> Metropolis-Hastings Correlation rho -> 0.6 Proposal width sigma -> 1 Seed -> 7 Burn-in -> 500

Watch the sample-mean, sample-variance and sample-cov(x, y) readouts settle onto the target, the acceptance and ESS(x) readouts, and the autocorrelation bars in the diagnostics panel. Every number you predict is printed there.

2 / 9
STAT 101
M9 · Lab
Step 1 of 5
Where the chain lives: the mean

Press Run and let the chain grow. The target's true mean is (0, 0). Predict the sample mean of x the chain converges to, then read the sample-mean readout.

Expected

The sample mean settles onto 0, the target's true mean. Because it is a sample average it jitters around the truth — after burn-in your readout sits within a few hundredths of (0, 0) — but it never drifts away: that settling is the first sign the chain has found the distribution.

3 / 9
STAT 101
M9 · Lab
Step 2 of 5
The spread: the variance

Leave every knob alone. The target's true variance is 1 in each coordinate. Predict the sample variance of x, then read the sample-variance readout.

Expected

The sample variance settles onto 1, the true marginal variance. Mean and variance together fix where the cloud sits and how wide it is; the last moment, the covariance, is the one the chain was never told.

4 / 9
STAT 101
M9 · Lab
Step 3 of 5
The headline: covariance → ρ

Still untouched. The target's off-diagonal covariance is ρ = 0.6 — the correlation you set, and the only thing that ties x and y together. Predict the sample cov(x, y) the chain converges to, then read the cov readout.

Expected

The sample cov(x, y) settles onto ρ = 0.6. This is the proof the sampler is honest: it reproduces the correlation structure of the target from nothing but accept/reject decisions on the density — the chain was never handed the covariance, it recovered it.

5 / 9
STAT 101
M9 · Lab
Step 4 of 5
Gibbs: an exact conditional

Switch Sampler to Gibbs. Instead of proposing and accepting, Gibbs draws each coordinate straight from its exact conditional, shown in the equations panel: x | y ~ N(ρy, 1 − ρ²). With ρ = 0.6, predict that conditional's standard deviation √(1 − ρ²), then confirm the acceptance readout shows n/a.

Expected

The conditional standard deviation is √(1 − ρ²) = √0.64 = 0.8, exactly — no sampling error, because it is the closed-form conditional the correlated normal happens to have. Every Gibbs step draws from it and is accepted by construction, so the acceptance readout shows n/a, and the sample cov still settles onto 0.6: same target, a different engine.

6 / 9
STAT 101
M9 · Lab
Step 5 of 5
Why n draws are worth fewer: acf[0]

Switch back to Metropolis-Hastings and look at the autocorrelation bars in the diagnostics panel. Predict the height of the very first bar, at lag 0, then read it.

Expected

The lag-0 autocorrelation is exactly 1 — any series is perfectly correlated with itself. It is the bars AFTER it that matter: the slower they decay, the more each draw repeats its neighbour, and the fewer effective independent samples ESS the chain is really worth.

7 / 9
STAT 101
M9 · Lab
Your turn
Open the sandbox

Everything above is waiting in the sandbox. Shrink the proposal width and watch acceptance shoot up while ESS collapses; widen it until the chain sticks; switch to the bimodal target and see a narrow proposal trap the chain in one mode and miss the other entirely — the classic mixing failure MCMC diagnostics exist to catch.

8 / 9
STAT 101
M9 · Lab
Wrap-up
What you did
  • Watched the sample mean and variance converge to the target's true (0, 0) and 1
  • Saw the sample cov(x, y) recover the correlation ρ = 0.6 — the proof the sampler is honest
  • Read Gibbs' exact conditional SD, √(1 − ρ²) = 0.8, every step accepted
  • Confirmed the lag-0 autocorrelation is exactly 1 — why ESS is below the chain length
  • Every value you predicted is the demo's own sampler arithmetic, not a picture
9 / 9