Reading
Stories Mode

The Bayesian Framework

~30 min read Lesson 1 of 2 in Module 9

Two Ways to Read a Probability

Everything in this course so far has been frequentist: a probability is the long-run frequency of an event, a parameter like a population mean is a fixed unknown constant, and a 95% confidence interval is a statement about the procedure, not about where the parameter is. That framework is powerful, but its central quantities — p-values, confidence intervals — are famously easy to misread, precisely because they do not answer the question people naturally ask: given my data, how likely is my hypothesis?

Bayesian statistics answers exactly that question, by treating probability as a degree of belief that can be updated by evidence. A parameter is not a fixed constant but a quantity we are uncertain about, described by a probability distribution. Data does not reveal the parameter; it updates our distribution over it. This one shift — from “the parameter is fixed and the data is random” to “the data is fixed and our belief is what changes” — is the whole of the Bayesian idea.

The Three Ingredients

Every Bayesian analysis is built from three distributions, and naming them is most of the battle:

Bayes’ Rule for Distributions

The three ingredients are tied together by Bayes’ theorem, applied not to single events but to whole distributions over θ:

Bayes’ Rule
P(\theta \mid D) = \dfrac{P(D \mid \theta)\,P(\theta)}{P(D)}
The posterior is the likelihood times the prior, divided by P(D) — the marginal likelihood, a normalising constant that makes the posterior integrate to one.

Because P(D) does not depend on θ, it only rescales the result. For most work the proportional form is all you need, and it reads as a sentence: the posterior is the prior, reweighted by how well each value of θ explains the data.

Posterior ∝ Likelihood × Prior
P(\theta \mid D) \;\propto\; P(D \mid \theta)\,P(\theta)
The engine of Bayesian updating. Strong data overwhelms a weak prior; a strong prior resists weak data; and with enough data the prior washes out entirely.

A Worked Example: Beta-Binomial

Suppose we want the probability θ that a coin (or a click, or a cure) succeeds. A natural prior for a probability is the Beta distribution, which lives on (0, 1) and is shaped by two counts, α and β:

Beta Prior
\theta \sim \text{Beta}(\alpha, \beta), \qquad p(\theta) \propto \theta^{\alpha-1}(1-\theta)^{\beta-1}
Read α−1 as “prior successes” and β−1 as “prior failures”. Beta(1, 1) is flat — total ignorance; Beta(50, 50) is a confident belief that θ is near one-half.

Now observe s successes and f failures. Multiply the Binomial likelihood by the Beta prior and the result is another Beta distribution, with the counts simply added on:

Conjugate Update
\text{Beta}(\alpha, \beta) \;\xrightarrow{\;s\text{ successes},\; f\text{ failures}\;}\; \text{Beta}(\alpha+s, \beta+f)
Updating is just arithmetic: add the observed successes to α and the observed failures to β. Each data point nudges the posterior; more data makes it narrower and more confident.

Conjugate Priors

The Beta-Binomial pair has a special property: the posterior is in the same family as the prior. A prior with this property is called a conjugate prior, and conjugate pairs (Beta-Binomial, Gamma-Poisson, Normal-Normal) turn Bayesian updating into closed-form algebra rather than a hard integral. They are the reason the framework was usable at all before computers — and they remain the clearest way to see the machinery, even though real problems usually need the computational methods this module’s deferred third lesson would cover.

Where the prior comes from is the framework’s most-debated feature: a critic calls it subjectivity, a practitioner calls it the honest inclusion of what was already known. Either way, the next lesson shows how to read the posterior once you have it — and how a Bayesian credible interval finally means what people always wanted a confidence interval to mean.

Key Takeaways
  • Frequentist vs. Bayesian: the frequentist treats the parameter as fixed and the data as random; the Bayesian treats the data as fixed and updates a probability distribution over the parameter.
  • Every analysis has three parts: the prior P(θ) (belief before data), the likelihood P(D∣θ), and the posterior P(θ∣D) (belief after data).
  • Bayes’ rule ties them together: posterior ∝ likelihood × prior. The marginal likelihood P(D) is just a normalising constant.
  • Strong data overwhelms a weak prior; with enough data the prior washes out. Updating reweights belief by how well each θ explains the data.
  • A conjugate prior keeps the posterior in the prior’s family — e.g. Beta-Binomial, where updating just adds successes to α and failures to β.
  • The choice of prior is the framework’s signature feature: a way to fold in what was already known, and the point critics scrutinise most.
Previous Non-Parametric Regression Module Overview Next Bayesian Estimation