Reading
Stories Mode

Bayes’ Theorem in Practice

~16 min read Lesson 3 of 3 in Module 3

When Mathematics Meets Real Decisions

A test for a rare disease comes back positive. A spam filter flags an incoming email. A product manager wants to know which version of a website is better. In all three cases, the question is the same: given some evidence, how should I update my belief?

Bayes’ theorem is the formal answer. In the previous lesson we derived it mathematically. Now we will use it — working through medical testing, spam filtering, and A/B experiments to see why Bayesian reasoning is both powerful and frequently violated by human intuition.

The Bayesian Recipe

Start with a prior belief P(H). Observe evidence E. Use the likelihood P(E|H) and the marginal probability P(E) to compute the posterior P(H|E).

Posterior ∝ Likelihood × Prior

Medical Testing: The Vocabulary

Medical tests are characterized by two fundamental quantities. Sensitivity is the probability that the test is positive given the patient truly has the disease — it measures how well the test catches true cases. Specificity is the probability that the test is negative given the patient is truly healthy — it measures how well the test avoids false alarms.

Sensitivity
\text{Sensitivity} = P(+ \mid D)
True positive rate: fraction of sick patients who test positive
Specificity
\text{Specificity} = P(- \mid \bar{D})
True negative rate: fraction of healthy patients who test negative

Notice what these quantities do not tell you. Sensitivity tells you P(positive | disease). But what a doctor and patient really want to know is the reverse: P(disease | positive). This is exactly where Bayes’ theorem comes in.

Positive Predictive Value and Negative Predictive Value

The quantities that answer clinical questions directly are the Positive Predictive Value (PPV) — the probability you have the disease given a positive test — and the Negative Predictive Value (NPV) — the probability you are healthy given a negative test. These are the posterior probabilities that Bayes’ theorem computes.

Positive Predictive Value (PPV)
PPV = \dfrac{P(+\mid D)\,P(D)}{P(+\mid D)\,P(D) + P(+\mid \bar{D})\,P(\bar{D})}
The denominator expands via the law of total probability, using prevalence P(D) and false-positive rate (1 − specificity)

The formula makes something immediately clear: PPV depends on prevalence P(D) — how common the disease is in the population being tested. A test with excellent sensitivity and specificity can still have a low PPV if the disease is rare, because most positive results come from the large pool of healthy people, not from the small pool of sick ones.

The Base Rate Fallacy

The base rate fallacy is the error of ignoring prevalence when interpreting test results. It is extraordinarily common, even among trained professionals. Studies have found that a majority of physicians significantly overestimate the probability of disease after a positive test, because they focus on the test’s accuracy and forget how rare the disease is.

Intuition tends to anchor on the headline number — “the test is 99% accurate” — and ignore the context — “but only 0.1% of people have the disease.” Bayes’ theorem forces us to weight the accuracy by the base rate, often with counterintuitive results.

Worked Example: Rare Disease Screening

Setup: A disease affects 1 in 1,000 people (prevalence = 0.1%). A test has 99% sensitivity and 99% specificity. You test positive. What is P(disease | positive)?

Numerator: P(positive | disease) × P(disease) = 0.99 × 0.001 = 0.00099

Denominator: P(positive) = 0.99 × 0.001 + 0.01 × 0.999 = 0.00099 + 0.00999 = 0.01098

PPV = 0.00099 / 0.01098 ≈ 9% — only a 9% chance of disease despite a “99% accurate” test

The result surprises almost everyone. Even with a highly accurate test, roughly 91 of 100 positive results come from healthy people, because there are 999 healthy people for every 1 sick person. The false positives from that large healthy pool swamp the true positives from the tiny sick pool.

Bayesian Updating: Prior → Evidence → Posterior

One of the most important properties of Bayes’ theorem is that it is sequential. Today’s posterior becomes tomorrow’s prior. You do not need to wait for all evidence before updating — you can incorporate each new piece of information one at a time, always using the most recent posterior as the starting point.

Suppose our patient from the example above undergoes a second, independent test that also comes back positive. We now take the posterior from the first test (PPV ≈ 9%) as the new prior, and apply Bayes’ theorem again with the second test’s likelihood. The posterior rises dramatically — because two independent positive tests are much harder to explain as coincidental false positives.

Sequential Bayesian Update
P(H \mid E_1, E_2) \propto P(E_2 \mid H)\,P(E_1 \mid H)\,P(H)
Each observation updates the posterior, which becomes the prior for the next observation — the likelihoods multiply when observations are independent

This sequential view is the heart of the Bayesian worldview: reasoning is an ongoing process of accumulating evidence and updating beliefs, never a one-shot calculation.

Spam Filtering with Naïve Bayes

Email spam filters were among the first large-scale deployments of Bayesian reasoning in software. The core idea: given the words in an email, what is the probability it is spam? We apply Bayes’ theorem treating each word as independent evidence — an approximation called the naïve Bayes assumption.

Let S = “email is spam” and let w1, w2, …, wn be the words in the email. With the naïve independence assumption, the posterior probability of spam is proportional to the prior probability of spam multiplied by the product of per-word likelihoods P(wi | S).

Naïve Bayes in Action

The filter is trained on labeled spam and ham emails, estimating P(word | spam) and P(word | ham) from counts. For a new email:

P(spam | words) ∝ P(spam) × ∏ P(wi | spam)

If this exceeds P(ham | words), the email is flagged. The “naïve” assumption ignores word co-occurrences, yet the classifier works remarkably well in practice despite this simplification.

Naïve Bayes is a generative classifier: it models how each class generates data. Despite the strong independence assumption, it is often competitive with much more complex models, especially when training data is limited. It is fast to train, interpretable, and naturally handles missing features.

Bayesian A/B Testing

In a classical A/B test, you run an experiment, compute a p-value, and decide whether to reject the null hypothesis. The Bayesian alternative starts with a prior distribution over each variant’s conversion rate and updates it with observed data to produce a posterior distribution. You can then directly ask: what is the probability that variant B is better than variant A?

The mechanism is the same prior → likelihood → posterior update from the medical-testing example, now applied to a continuous conversion rate θ rather than a yes/no diagnosis. You begin with a prior belief about θ — a whole distribution over its plausible values, not a single guess — and every visitor who does or does not convert is one more observation that sharpens it. As conversions accumulate, the posterior concentrates around the true rate. The specific distributions that make this update exact and closed-form are introduced in Module 4; here the point is only that Bayesian A/B testing is the same sequential update, run over and over.

Bayesian vs. Frequentist A/B

Frequentist: “The p-value is 0.03, so we reject H₀ at α = 0.05.” Cannot say probability that B is better.

Bayesian: “P(B better than A | data) = 94%.” Direct answer to the business question, with a credible interval on the lift.

Bayesian tests also make continuous monitoring easier to reason about: a posterior is interpretable at any sample size, so looking at the data early does not invalidate the posterior itself. But this is not a free pass — a decision rule such as “stop as soon as P(B better) exceeds 95%” is still a stopping rule, and it still inflates the long-run false-positive rate. Bayesian methods change what the number means, not the arithmetic of repeated looks.

Why Bayes Matters for Machine Learning

Bayesian ideas permeate modern machine learning. Maximum a posteriori (MAP) estimation is Bayes’ theorem applied to parameter estimation: instead of maximizing the likelihood alone (maximum likelihood estimation), MAP maximizes the product of likelihood and prior, which is equivalent to regularization. L2 regularization (ridge) corresponds to a Gaussian prior; L1 regularization (lasso) corresponds to a Laplace prior.

Bayesian neural networks maintain a full posterior distribution over weights rather than a point estimate, enabling principled uncertainty quantification. Gaussian processes are Bayesian models over functions, used in hyperparameter optimization and spatial statistics. Variational inference and MCMC methods make these computations tractable.

The Bayesian Mindset in ML

Uncertainty is first-class. A model that knows what it does not know is more useful than one that is always confident.

Prior knowledge is data. Domain expertise enters naturally as a prior, allowing models to generalize better when labeled data is scarce.

Learning is updating. Every new observation is an opportunity to refine beliefs — the same principle scales from coin flips to training deep networks.

Putting It All Together

Bayes’ theorem is not just a formula — it is a framework for rational belief revision. In medical testing, it shows that a positive result from a rare-disease screen is almost certainly a false positive unless the prior probability of disease is substantial. In spam filtering, it powers classifiers that adapt to individual users’ email patterns. In A/B testing, it provides direct answers to business questions rather than indirect p-value thresholds. And in machine learning, it connects regularization, uncertainty quantification, and sequential learning under a single coherent umbrella.

The base rate fallacy — ignoring prevalence — is perhaps the most important lesson of this module. Whenever you see a diagnostic test result, a model confidence score, or a striking experimental finding, ask: what was the prior? The answer will often change everything.

Key Takeaways
  • Sensitivity = P(positive | disease); specificity = P(negative | healthy). Neither tells you P(disease | positive) — that requires Bayes.
  • PPV = P(disease | positive) depends critically on prevalence. A “99% accurate” test for a 0.1%-prevalent disease has PPV ≈ 9%.
  • The base rate fallacy is ignoring the prior (prevalence). Bayes’ theorem forces prior and likelihood to combine correctly.
  • Bayesian updating is sequential: today’s posterior is tomorrow’s prior. Independent observations multiply their likelihoods.
  • Naïve Bayes spam filters apply Bayes’ theorem word-by-word under an independence assumption, yet work surprisingly well.
  • Bayesian A/B testing gives P(B better than A | data) directly and makes mid-experiment monitoring interpretable — but the stopping rule still has to be chosen deliberately; it does not make repeated looks free.
  • In ML, Bayes appears as MAP estimation (regularization), Bayesian neural networks, Gaussian processes, and principled uncertainty quantification.
Previous Conditional Probability Module Overview Next Module Random Variables