The Posterior Is the Whole Answer
The previous lesson produced a posterior distribution P(θ ∣ D) — a full probability distribution over the parameter, given the data. That distribution is the Bayesian answer: it already contains everything we know about θ, including our uncertainty. Estimation, in the Bayesian world, is not about finding a single “true” value but about summarising that posterior for a human who has to make a decision.
Point Estimates from the Posterior
When a single number is needed, the posterior offers three natural ones. The posterior mean is the expected value of θ under the posterior — the most common Bayesian point estimate because it minimises expected squared error:
The posterior mode — the most probable single value — is called the maximum a posteriori (MAP) estimate. With a flat prior it coincides exactly with the frequentist maximum-likelihood estimate, which is the cleanest bridge between the two worlds:
The Credible Interval
For an interval estimate, the Bayesian reports a credible interval: a range that contains a stated fraction — usually 95% — of the posterior probability.
Credible vs. Confidence — The Interval People Wanted
This is the payoff. In Module 5 we were careful — and emphatic — that a 95% confidence interval does not mean “there is a 95% probability the parameter is in this interval.” In the frequentist world the parameter is fixed, so it is either in a given interval or not; the 95% describes the long-run behaviour of the procedure, not this interval. That caveat frustrates every student, because the forbidden reading is exactly the one everyone wants.
The Bayesian credible interval means precisely the forbidden thing. Because θ has a probability distribution, the statement “there is a 95% probability that θ is between θL and θU” is not only allowed — it is the definition. The interval that Module 5 warned you not to misread is the one this framework hands you honestly. That single difference — a probability statement about the parameter rather than about the procedure — is the practical reason many fields have moved toward Bayesian reporting.
Choosing a Prior
An uninformative (flat) prior expresses deliberate ignorance and lets the data dominate; with it, Bayesian and frequentist point estimates often coincide. An informative prior folds in genuine prior knowledge — a physical constraint, an earlier study — and pulls the posterior toward it, which is valuable when data is scarce and dangerous when the prior is wrong. Because a strong prior can sway a conclusion, honest Bayesian practice includes a sensitivity analysis: re-run the analysis under several reasonable priors and report how much the conclusion moves. If it barely moves, the data is speaking; if it swings, the prior is, and your reader deserves to know.
Bayesian Linear Regression
Everything composes with the regression of Module 7. In Bayesian linear regression, the coefficients β get priors and the output is a posterior distribution over each coefficient — so instead of a point estimate and a confidence interval you get a full credible interval per coefficient, and predictions arrive with honest uncertainty bands. A Gaussian prior on β reproduces Ridge regression as its MAP estimate, and a Laplace prior reproduces Lasso: the regularisation of Module 7 and the priors of Module 9 are two views of one idea.
- The posterior distribution is the full Bayesian answer; estimation is summarising it for a decision.
- Point estimates: posterior mean (minimises squared error), median (robust), and MAP (the mode). Under a flat prior, MAP equals the maximum-likelihood estimate.
- A credible interval contains a stated fraction of the posterior probability — e.g. a 95% probability that θ is inside it.
- Credible ≠ confidence: the credible interval means exactly what Module 5 forbade you to read into a confidence interval — a probability statement about the parameter, not the procedure.
- Priors range from flat (let the data speak) to informative (fold in knowledge); a sensitivity analysis reports how much the prior is driving the conclusion.
- Bayesian linear regression gives a posterior per coefficient; Gaussian and Laplace priors reproduce Ridge and Lasso as MAP estimates.
The module closes with 9.3 — Computational Bayesian Methods (MCMC): Metropolis-Hastings, Gibbs sampling, and diagnostics. Those methods are hard to follow as static prose, so the lesson pairs with an interactive sampler where you watch a chain converge and mix as you turn the proposal width. The “Next” button goes there.