Reading
Stories Mode

Bayesian Estimation

~30 min read Lesson 2 of 2 in Module 9

The Posterior Is the Whole Answer

The previous lesson produced a posterior distribution P(θ ∣ D) — a full probability distribution over the parameter, given the data. That distribution is the Bayesian answer: it already contains everything we know about θ, including our uncertainty. Estimation, in the Bayesian world, is not about finding a single “true” value but about summarising that posterior for a human who has to make a decision.

Point Estimates from the Posterior

When a single number is needed, the posterior offers three natural ones. The posterior mean is the expected value of θ under the posterior — the most common Bayesian point estimate because it minimises expected squared error:

Posterior Mean
\hat{\theta}_{\text{Bayes}} = \mathbb{E}[\theta \mid D] = \int \theta\, p(\theta \mid D)\, d\theta
The centre of mass of the posterior. The posterior median (robust, minimises absolute error) is an alternative when the posterior is skewed.

The posterior mode — the most probable single value — is called the maximum a posteriori (MAP) estimate. With a flat prior it coincides exactly with the frequentist maximum-likelihood estimate, which is the cleanest bridge between the two worlds:

MAP Estimate
\hat{\theta}_{\text{MAP}} = \arg\max_{\theta}\, p(\theta \mid D)
The peak of the posterior. Under a uniform prior, MAP = the maximum-likelihood estimate — regularised regression’s penalties are exactly MAP estimates under Gaussian (Ridge) or Laplace (Lasso) priors.

The Credible Interval

For an interval estimate, the Bayesian reports a credible interval: a range that contains a stated fraction — usually 95% — of the posterior probability.

95% Credible Interval
P(\theta_L \le \theta \le \theta_U \mid D) = 0.95
There is a 95% posterior probability that θ lies between θL and θU. The highest-density version reports the shortest such interval.

Credible vs. Confidence — The Interval People Wanted

This is the payoff. In Module 5 we were careful — and emphatic — that a 95% confidence interval does not mean “there is a 95% probability the parameter is in this interval.” In the frequentist world the parameter is fixed, so it is either in a given interval or not; the 95% describes the long-run behaviour of the procedure, not this interval. That caveat frustrates every student, because the forbidden reading is exactly the one everyone wants.

The Bayesian credible interval means precisely the forbidden thing. Because θ has a probability distribution, the statement “there is a 95% probability that θ is between θL and θU” is not only allowed — it is the definition. The interval that Module 5 warned you not to misread is the one this framework hands you honestly. That single difference — a probability statement about the parameter rather than about the procedure — is the practical reason many fields have moved toward Bayesian reporting.

Choosing a Prior

An uninformative (flat) prior expresses deliberate ignorance and lets the data dominate; with it, Bayesian and frequentist point estimates often coincide. An informative prior folds in genuine prior knowledge — a physical constraint, an earlier study — and pulls the posterior toward it, which is valuable when data is scarce and dangerous when the prior is wrong. Because a strong prior can sway a conclusion, honest Bayesian practice includes a sensitivity analysis: re-run the analysis under several reasonable priors and report how much the conclusion moves. If it barely moves, the data is speaking; if it swings, the prior is, and your reader deserves to know.

Bayesian Linear Regression

Everything composes with the regression of Module 7. In Bayesian linear regression, the coefficients β get priors and the output is a posterior distribution over each coefficient — so instead of a point estimate and a confidence interval you get a full credible interval per coefficient, and predictions arrive with honest uncertainty bands. A Gaussian prior on β reproduces Ridge regression as its MAP estimate, and a Laplace prior reproduces Lasso: the regularisation of Module 7 and the priors of Module 9 are two views of one idea.

Key Takeaways
  • The posterior distribution is the full Bayesian answer; estimation is summarising it for a decision.
  • Point estimates: posterior mean (minimises squared error), median (robust), and MAP (the mode). Under a flat prior, MAP equals the maximum-likelihood estimate.
  • A credible interval contains a stated fraction of the posterior probability — e.g. a 95% probability that θ is inside it.
  • Credible ≠ confidence: the credible interval means exactly what Module 5 forbade you to read into a confidence interval — a probability statement about the parameter, not the procedure.
  • Priors range from flat (let the data speak) to informative (fold in knowledge); a sensitivity analysis reports how much the prior is driving the conclusion.
  • Bayesian linear regression gives a posterior per coefficient; Gaussian and Laplace priors reproduce Ridge and Lasso as MAP estimates.
Next: Computational Bayesian Methods

The module closes with 9.3 — Computational Bayesian Methods (MCMC): Metropolis-Hastings, Gibbs sampling, and diagnostics. Those methods are hard to follow as static prose, so the lesson pairs with an interactive sampler where you watch a chain converge and mix as you turn the proposal width. The “Next” button goes there.

Previous The Bayesian Framework Module Overview Next Computational Bayesian Methods